Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

6,963 resources indexed

Showing 6,7016,750

The tool is designed to perform a customisable image pre-processing to reduce noise and inhomogeneity field effect, thus improving image quality and reproducibility of radiomics features. This tool consists of two independent steps: one for denoising using one of the 5 integrated filters (Bilateral Filter, Anisotropic Diffusion Filter (ADF), Curvature Flow Filter (CFF), SUSAN and Non Local Means (NLM)), and another for the ANTs N4 and another for the ANT's N4 bias correction filter. The parameter configuration of this tool has been optimised for TW1, T2W, DWI and DCE sequences in neuroblastoma (NB) and paediatric brain tumours, but it can also be configured with some of their parameters using a JSON parameter configuration file.

A tool based on artificial intelligence that is able to perform a categorisation of MRI series by using standardized DICOM tags. The categorisation includes the type of sequence (e.g. spin echo, gradient echo), the weighting (e.g. T1W, T2W, DCE, ...), the presence of fat suppression and the detection of non-relevant / junk series (e.g. localizers, calibrations, screenshots...).

Tool that aims to validate visually the chronological order and logical consistency of dates associated with a patient's medical history. It generates a timeline visualization for each patient from an Excel file and highlights rule violations. Status : Containerized

The tool performs a DICOM quality check in terms of correct number of files per sequence, corrupted files, precise directory hierarchy, separated dynamic series merging them, interest series filtering/selection by specific series description lists and diffusion sequence identification by b-values. It applies the desired changes to the dataset and generates a report containing information about the selected sequences, corrupted files, missing files and merged files. Status: Deployed

VIP is a web portal for medical imaging applications. It allows users to access scientific applications as a service (directly through the web browser with no installation required), as well as distributed computing resources in a transparent manner. It exploits the resources available in the biomed virtual organization of the EGI e-infrastructure to offer an open service to researchers worldwide.

Membrane Protein-Lipid Interaction Database. A large-scale experimentally validated dataset of 80685 residue-level lipid contact annotations across 4712 membrane proteins derived from PDB crystal and cryo-EM structures. Provides pre-computed binary contact labels, continuous distance values, sequence-identity-based cluster assignments, and ready-made train-validation-test splits for machine learning.

ekokrati computes habitat connectivity metrics (PC, IIC, EC(PC), dPC and its decomposition into intra-patch, flux and connector components) for habitat patch networks. Users upload polygon data as GeoPackage or shapefile, set species-specific dispersal parameters, and receive patch importance scores and landscape-level indices. Designed for conservation planners, landscape ecologists and environmental consultants. No installation required.

AI scientist framework for autonomous deep research in biological sciences, combining literature analysis agents with data scientist agents to enable iterative scientific discovery through user feedback integration; achieves state-of-the-art performance on BixBench benchmark (48.78% open-answer, 64.39% multiple-choice) outperforming Kepler and GPT-5 (bio-xyz, arXiv 2601.12542, 160+ stars, 2025-2026)

Verbex is a private, on-device Voice-to-ELN iOS app for scientists. It helps researchers capture experiment notes by voice as work happens, organize those notes into scientific sections, and prepare clean, reviewable, ELN-ready scientific records.

The MetaProteomeAnalyzer Cloud (MPA Cloud) is an intuitive, open-source tool for metaproteomics data analysis and interpretation, designed to analyse comprehensive metaproteomics data from tandem mass spectrometry experiments through a web interface.

Open Babel is a chemical toolbox designed to speak the many languages of chemical data.

Molecular Biology Tools is a free browser-based collection of molecular biology utilities for routine sequence analysis, primer design, Sanger sequencing primer planning, cloning setup, and wet-lab calculations. The site includes tools for PCR primer design, Sanger primer design and primer walking, primer binding checks, restriction site analysis, reverse complement generation, ORF and protein translation, codon optimization, ligation calculations, molarity calculations, dilution calculations, and multi-solute solution recipe preparation. The tools run in the browser and are intended for quick experimental planning, without requiring logins or uploading sequences to the server.

Comprehensive set of programs for phylogenetic analyses; available for PC and Mac; source code available for easy compiling in UNIX.

FlavoTyper is a bioinformatics tool that performs in silico serotyping of Flavobacterium psychrophilum genome assemblies.

MONAI Label is an intelligent open source image labeling and learning tool that enables users to create annotated datasets and build AI annotation models for clinical evaluation. MONAI Label enables application developers to build labeling apps in a serverless way, where custom labeling apps are exposed as a service through the MONAI Label Server.

The ProteinsPlus web server aims to support life scientists in working with protein structures. Protein structures are the key to understanding protein function. They are an important resource in many biotechnological application areas from pharmaceutical research to biocatalysis. ProteinsPlus focuses on protein-ligand interactions. The server provides support for the initial steps of dealing with protein structures, namely structure search, quality assessment, and preprocessing. JAMDA enables users to perform an on-the-fly molecular docking of up to five molecules. The poses can then be visualized in 2D (PoseView, PoseEdit). Furthermore, advanced options, such as protein pocket detection (DoGSite), prediction of water molecule positions (WarPP), protein structure ensemble generation (SIENA), prediction of metal coordination (METALizer), the analysis of solvent channels in protein crystals (LifeSoaks), or the categorization of protein-protein-interfaces (HyPPI) are supported.

SeqBench is a browser-based workbench of 82 tools for molecular cloning, primer design and sequence analysis: nearest-neighbour melting temperature, oligo dimer and hairpin screening, in-silico PCR, site-directed mutagenesis, restriction mapping, Gibson, Golden Gate and restriction-ligation assembly simulation, plasmid annotation and backbone identification, CRISPR gRNA, HDR donor and base editing design, codon optimisation and CAI, pairwise and multiple alignment, RNA secondary structure, protein properties, Sanger ab1 traces, HGVS conversion and variant annotation. It verifies constructs as well as designing them: re-deriving an assembly from its stated parts and diffing it against the claimed product, aligning sequencing reads back onto a claimed reference, and scoring Golden Gate overhang sets against published ligation-fidelity data. Batch FASTA processing, multi-tool workflows, a conversational tool-calling interface (SeqBench-GPT), a REST API and an MCP server are included.

regenie is a C++ program for whole genome regression modelling of large genome-wide association studies.

Bin Chicken - recovery of low abundance and taxonomically targeted metagenome assembled genomes (MAGs) through strategic coassembly

Community Terrestrial Systems Model (includes the Community Land Model of CESM)

Thoa is a cloud bioinformatics platform. Write your Nextflow or Snakemake pipeline, point it at your data, and Thoa handles the rest: provisioning VMs (up to 12TB RAM), resolving dependencies, managing execution. No cloud expertise needed. Every job captures its full context:data, software versions, environment, machine specs, as a reproducibility artifact. Share it with a colleague and they can view or re-run the analysis without an account. Key features: AI debugger that fixes environment and dependency issues in real time. Pipeline tracking with per-step telemetry. if step 47 of 200 fails, re-run from there, not from scratch. One-click data sharing without registration. AI-assisted workflow creation from plain English. Free tier available. Starter $35/mo, Pro $109/mo, Team $480/mo. Zero-egress storage. Based in Zug, Switzerland. thoa.io​​​​​​​​​​​​​​​​

URGI (Unit Resources Genomics-Info) is a bioinformatics facility who support various research activities on plants of agronomic and forestry interest for INRAE. The platform has federated with 3 other INRAE ​​bioinformatics platforms to form the BioinfOmics research infrastructure. It is part of the French Institute of Bioinformatics which is the French node of the European infrastructure ELIXIR . URGI is part of the Saclay Plant Science network and of the Graduate School Biosphera . URGI is labeled by GIS IBiSA and is ISO-9001 certified.

Comparative genomics-driven translational research tool.

metagWGS is a workflow dedicated to the analysis of metagenomic data. It allows assembly, taxonomic annotation, and functional annotation of predicted genes. Since release 2.3, binning step with the possibility of cross-alignment is included. It has been developed in collaboration with several CATI BIOS4biol agents. Funded by Antiselfish Project (Labex Ecofect), ExpoMicoPig project (France Futur elevage) and SeqOccIn project (CPER - Occitanie Toulouse / FEDER), ATB_Biofilm funded by PNREST Anses, France genomique (ANR-10-INBS-09-08) and Resalab Ouest.

Browser-based viewer for Sanger sequencing chromatograms in AB1/ABIF and SCF format, and for the .srd raw files of the Nanofor-05 capillary sequencer, converted to ABIF as they open. Shows raw and analysed traces, base calls (called or edited), per-base quality and the full ABIF directory, and puts two reads side by side with their tag tables aligned. A read aligns against a pasted or loaded reference, which reports identity, mismatches, indels and the ends that did not align, and flips the strand when the read is on the other one. End trimming (modified Mott or sliding window) has draggable handles, motif search is IUPAC-aware and finds primers, and QC metrics copy out as CSV. Drag on a trace to select a base range, then copy it as FASTA, zoom to it or export just that region. Exports the read as FASTA, FASTQ, .qual or ABIF and the chromatogram as a high-resolution PNG. Open files are kept as local sessions. Runs entirely in the browser — files are never uploaded.

Conspecta is a browser-based research platform that brings microscopy image analysis, flow cytometry, molecular biology, sample tracking, and publication-ready figures into one connected workspace. It replaces the patchwork of disconnected tools most labs assemble, so a lab's data, samples, and results stay linked from experiment to figure with full traceability. Built for imaging-heavy and flow-heavy biology labs, new PIs, and early-stage biotech. Research-focused, not regulated or clinical. Free for you and one collaborator, every workspace included. Paid plans open it to the whole lab and add additional storage, AI object detection, external integrations via API, and bring-your-own AI assistant integration.

xnattools is a python package with a set of modules for performing various operations on data stored in XNAT servers. The main purpose is to provide one standardized platform for running operations on XNAT servers. The package currently contains four tools that use this platform: dicom to nifty conversion, thumbnail generation from dicom, DICOM header data collection, in bulk downloading of a project.

BioCalculator is a mobile toolkit for biology and laboratory calculations. It consolidates calculators and preparation guides for molecular biology, microbiology, virology, stock solutions, buffers, dilutions, unit conversions, and TCID50 endpoint dilution workflows for students and lab professionals.

Python package for biodatafuse project.

AmsterdamUMCdb is a database of de-identified health data related to tens of thousands of intensive care unit admissions, including demographics, vital signs, laboratory tests and medications.

austraits is an R package for accessing the AusTraits Plant database and working with traits.build databases

Browser-based tool to open almost any file that carries sequence — FASTA, FASTQ, GenBank, EMBL, Swiss-Prot, AB1/ABIF, SCF, Clustal, Stockholm, PHYLIP, NEXUS, MSF, PIR, MEGA, GFF3, SAM, BAM, GFA, PDB and ACE — see every sequence inside, pick the ones you want, and save them as FASTA. The format is detected from the file content, not from the extension, so unlabelled or misnamed files still open, and gzip-compressed files are unpacked in place. Records can be filtered by length, name, GC or sequence type, reverse-complemented, transcribed DNA↔RNA or translated to protein, deduplicated and reordered before saving. Runs entirely in the browser — files are never uploaded.

VigyanLLM is a sovereign, on-premises biomedical AI platform designed for computational biology. It provides autonomous primer design, evaluating melting temperature (Tm) and GC content; CRISPR guide RNA analysis with off-target scoring; BLAST sequence similarity searching; multiple sequence alignment (MSA); and GPU-accelerated molecular docking for drug discovery. Unlike cloud-based SaaS, VigyanLLM deploys via Docker to ensure institutional genomic data sovereignty

AI Agent for Biomedical Research

SciAgentKit is an MCP-native toolkit that connects AI agents to reproducible computational drug-discovery workflows. It integrates established tools for molecular analysis, protein-structure assessment, binding-site detection, molecular docking, molecular dynamics, trajectory analysis and scientific reporting.

Trace4MedicalImageCleaning™ is a tool aimed at automatically detecting and removing text in medical images, with a specific focus on 2D ultrasound and mammography studies.

NIM Studio is a local-first platform for neuroinformatics, BIDS organization, metadata generation, duplicate auditing, and scalable research data management.

A bioinformatics tool for viewing and calculating base modification frequencies from BAM files

Simple test framework for Nextflow pipelines

Reactr is an modularized, Snakemake workflow for automated, species-agnostic characterization of gene families from sequence to experimental design. Given a query protein sequence and NCBI taxonomy IDs (or RefSeq assembly accessions), reactr retrieves genomic data and runs comprehensive analysis across 4 integrated tiers: (1) evolutionary analysis, including homolog detection, domain-based clustering, multiple sequence alignment, and phylogenetic inference; (2) synteny and selection analysis, detecting collinear blocks and calculating Ka/Ks ratios; (3) structural and regulatory characterization, including motif discovery, chromosomal mapping, biochemical property prediction, subcellular localization prediction, and promoter analysis; and (4) experimental design tools, generating PCR primers and scored CRISPR gRNAs for lab validation. Reactr bridges computational prediction and experimental validation, thus enabling rapid transition from genomic discovery to functional studies.

Harmonize numerical values extracted from medical images (e.g. acquired with different models of image-acquisition system)

Browser-based viewer for GenBank and GenPept records — .gb, .gbk, .gbff, .gp and plain GenBank text. Renders an interactive linear and circular feature map, including circular plasmid maps, alongside the annotated source text and the nucleotide/protein sequence. Translates CDS features using the record's own genetic code and translation qualifiers, flags where the stored /translation disagrees with a plain translation, and adds optional computed layers: ORF prediction and restriction-site mapping. Drag a range in the sequence band to select it, then copy that stretch — either strand, as DNA or as protein, plain or FASTA. Filters features by type, handles multi-record files, and keeps open records as local sessions, so a closed tab can be picked up where it was left. Runs entirely in the browser — files are never uploaded.

ArrayAnalysis is a web-based application for transcriptomic data analysis. It supports the analysis of both microarray and RNA-seq data. The tool may also be installed locally as a desktop app, Docker image, or R package.

Supernova is a software package for de novo assembly from Chromium Linked-Reads that are made from a single whole-genome library from an individual DNA source. A key feature of Supernova is that it creates diploid assemblies, thus separately representing maternal and paternal chromosomes over very long distances. Almost all other methods instead merge homologous chromosomes into single incorrect 'consensus' sequences. Supernova is the only practical method for creating diploid assemblies of large genomes.

edf2csv is a local command-line tool for converting EDF, EDF+, BDF, and BDF+ physiological recordings into CSV and JSON files. It exports signal values, channel information, annotations, and recording metadata while preserving original sampling rates, physical units, and discontinuities.

Plotwright is a browser-based statistics and scientific graphing workbench for life-science researchers. It provides a Data to Analysis to Graph workflow for statistical tests, regression and curve fitting, survival and ROC analysis, and editable scientific figures. It can import supported content from modern GraphPad Prism .prism files and reports preserved, approximated, and unsupported objects for review. Projects can be saved as local files. Plotwright is commercial software with a 14-day no-card trial.

Predicting the effect of mutations on protein-RNA binding with Deep Learning | This repository contains all DeepCLIP Python code | A context-aware neural network for modeling and predicting protein binding to nucleic acids using only sequence input | DeepCLIP is a neural network with shallow convolutional layers connected to a bidirectional LSTM layer

Plant Compound Extractor is a desktop application that builds a ready-to-use, deduplicated library of 3D ligand structures for a given plant. It queries multiple natural-product and chemical databases (COCONUT, LOTUS, Wikidata, PubChem, PlantaeDB, USDA Dr. Duke's, KNApSAcK and IMPPAT) in parallel, resolves each compound against PubChem for a canonical structure, and falls back to direct source retrieval when needed. Retrieved structures are then converted to 3D using RDKit, with configurable conformer generation and physicochemical filters (molecular weight, rotatable bonds, ring size, etc.). It can also process a manually supplied compound list, or convert an existing folder of 2D structures to 3D.

Desktop viewer for microscopy and whole slide pathology images on Windows and macOS. Opens whole slide scanner formats (Aperio SVS, Hamamatsu NDPI, MIRAX MRXS, Leica SCN, Ventana BIF) alongside acquisition formats (Zeiss CZI, Nikon ND2, DICOM, TIFF/OME-TIFF) in a single application, providing pyramid navigation of gigapixel images, metadata inspection, calibrated measurements, annotations, and on-device segmentation and object counting.

Pan.bio is a cloud genomics platform for pipeline execution, exploratory analysis, and clinical variant interpretation. Workflows runs validated Nextflow and nf-core pipelines including Sarek, rnaseq, scrnaseq, mag, ampliseq, chipseq and atacseq without local installation. Notebooks provides Python and R sessions with a preinstalled bioinformatics stack, importing public data from GEO, SRA and IPG by accession and reading Workflows outputs directly. VAIC applies ACMG/AMP variant classification with gene-specific rule sets from CanVIG-UK and ClinGen ENIGMA, with automated evidence criteria implemented for BRCA1 and BRCA2. Cohorts provides federated analysis of patient data within a Trusted Research Environment.