Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

291 of 6,573 resources

Showing 251291

SeqBench is a browser-based workbench of 82 tools for molecular cloning, primer design and sequence analysis: nearest-neighbour melting temperature, oligo dimer and hairpin screening, in-silico PCR, site-directed mutagenesis, restriction mapping, Gibson, Golden Gate and restriction-ligation assembly simulation, plasmid annotation and backbone identification, CRISPR gRNA, HDR donor and base editing design, codon optimisation and CAI, pairwise and multiple alignment, RNA secondary structure, protein properties, Sanger ab1 traces, HGVS conversion and variant annotation. It verifies constructs as well as designing them: re-deriving an assembly from its stated parts and diffing it against the claimed product, aligning sequencing reads back onto a claimed reference, and scoring Golden Gate overhang sets against published ligation-fidelity data. Batch FASTA processing, multi-tool workflows, a conversational tool-calling interface (SeqBench-GPT), a REST API and an MCP server are included.

regenie is a C++ program for whole genome regression modelling of large genome-wide association studies.

RxDock is a fast and versatile open-source docking program that can be used to dock small molecules against proteins and nucleic acids. It is designed for high-throughput virtual screening (HTVS) campaigns and binding mode prediction studies.

Bin Chicken - recovery of low abundance and taxonomically targeted metagenome assembled genomes (MAGs) through strategic coassembly

Community Terrestrial Systems Model (includes the Community Land Model of CESM)

Thoa is a cloud bioinformatics platform. Write your Nextflow or Snakemake pipeline, point it at your data, and Thoa handles the rest: provisioning VMs (up to 12TB RAM), resolving dependencies, managing execution. No cloud expertise needed. Every job captures its full context:data, software versions, environment, machine specs, as a reproducibility artifact. Share it with a colleague and they can view or re-run the analysis without an account. Key features: AI debugger that fixes environment and dependency issues in real time. Pipeline tracking with per-step telemetry. if step 47 of 200 fails, re-run from there, not from scratch. One-click data sharing without registration. AI-assisted workflow creation from plain English. Free tier available. Starter $35/mo, Pro $109/mo, Team $480/mo. Zero-egress storage. Based in Zug, Switzerland. thoa.io​​​​​​​​​​​​​​​​

URGI (Unit Resources Genomics-Info) is a bioinformatics facility who support various research activities on plants of agronomic and forestry interest for INRAE. The platform has federated with 3 other INRAE ​​bioinformatics platforms to form the BioinfOmics research infrastructure. It is part of the French Institute of Bioinformatics which is the French node of the European infrastructure ELIXIR . URGI is part of the Saclay Plant Science network and of the Graduate School Biosphera . URGI is labeled by GIS IBiSA and is ISO-9001 certified.

Comparative genomics-driven translational research tool.

metagWGS is a workflow dedicated to the analysis of metagenomic data. It allows assembly, taxonomic annotation, and functional annotation of predicted genes. Since release 2.3, binning step with the possibility of cross-alignment is included. It has been developed in collaboration with several CATI BIOS4biol agents. Funded by Antiselfish Project (Labex Ecofect), ExpoMicoPig project (France Futur elevage) and SeqOccIn project (CPER - Occitanie Toulouse / FEDER), ATB_Biofilm funded by PNREST Anses, France genomique (ANR-10-INBS-09-08) and Resalab Ouest.

Browser-based viewer for Sanger sequencing chromatograms in AB1/ABIF and SCF format, and for the .srd raw files of the Nanofor-05 capillary sequencer, converted to ABIF as they open. Shows raw and analysed traces, base calls (called or edited), per-base quality and the full ABIF directory, and puts two reads side by side with their tag tables aligned. A read aligns against a pasted or loaded reference, which reports identity, mismatches, indels and the ends that did not align, and flips the strand when the read is on the other one. End trimming (modified Mott or sliding window) has draggable handles, motif search is IUPAC-aware and finds primers, and QC metrics copy out as CSV. Drag on a trace to select a base range, then copy it as FASTA, zoom to it or export just that region. Exports the read as FASTA, FASTQ, .qual or ABIF and the chromatogram as a high-resolution PNG. Open files are kept as local sessions. Runs entirely in the browser — files are never uploaded.

Conspecta is a browser-based research platform that brings microscopy image analysis, flow cytometry, molecular biology, sample tracking, and publication-ready figures into one connected workspace. It replaces the patchwork of disconnected tools most labs assemble, so a lab's data, samples, and results stay linked from experiment to figure with full traceability. Built for imaging-heavy and flow-heavy biology labs, new PIs, and early-stage biotech. Research-focused, not regulated or clinical. Free for you and one collaborator, every workspace included. Paid plans open it to the whole lab and add additional storage, AI object detection, external integrations via API, and bring-your-own AI assistant integration.

xnattools is a python package with a set of modules for performing various operations on data stored in XNAT servers. The main purpose is to provide one standardized platform for running operations on XNAT servers. The package currently contains four tools that use this platform: dicom to nifty conversion, thumbnail generation from dicom, DICOM header data collection, in bulk downloading of a project.

BioCalculator is a mobile toolkit for biology and laboratory calculations. It consolidates calculators and preparation guides for molecular biology, microbiology, virology, stock solutions, buffers, dilutions, unit conversions, and TCID50 endpoint dilution workflows for students and lab professionals.

Python package for biodatafuse project.

AmsterdamUMCdb is a database of de-identified health data related to tens of thousands of intensive care unit admissions, including demographics, vital signs, laboratory tests and medications.

austraits is an R package for accessing the AusTraits Plant database and working with traits.build databases

Browser-based tool to open almost any file that carries sequence — FASTA, FASTQ, GenBank, EMBL, Swiss-Prot, AB1/ABIF, SCF, Clustal, Stockholm, PHYLIP, NEXUS, MSF, PIR, MEGA, GFF3, SAM, BAM, GFA, PDB and ACE — see every sequence inside, pick the ones you want, and save them as FASTA. The format is detected from the file content, not from the extension, so unlabelled or misnamed files still open, and gzip-compressed files are unpacked in place. Records can be filtered by length, name, GC or sequence type, reverse-complemented, transcribed DNA↔RNA or translated to protein, deduplicated and reordered before saving. Runs entirely in the browser — files are never uploaded.

VigyanLLM is a sovereign, on-premises biomedical AI platform designed for computational biology. It provides autonomous primer design, evaluating melting temperature (Tm) and GC content; CRISPR guide RNA analysis with off-target scoring; BLAST sequence similarity searching; multiple sequence alignment (MSA); and GPU-accelerated molecular docking for drug discovery. Unlike cloud-based SaaS, VigyanLLM deploys via Docker to ensure institutional genomic data sovereignty

AI Agent for Biomedical Research

SciAgentKit is an MCP-native toolkit that connects AI agents to reproducible computational drug-discovery workflows. It integrates established tools for molecular analysis, protein-structure assessment, binding-site detection, molecular docking, molecular dynamics, trajectory analysis and scientific reporting.

Trace4MedicalImageCleaning™ is a tool aimed at automatically detecting and removing text in medical images, with a specific focus on 2D ultrasound and mammography studies.

NIM Studio is a local-first platform for neuroinformatics, BIDS organization, metadata generation, duplicate auditing, and scalable research data management.

A bioinformatics tool for viewing and calculating base modification frequencies from BAM files

Simple test framework for Nextflow pipelines

Reactr is an modularized, Snakemake workflow for automated, species-agnostic characterization of gene families from sequence to experimental design. Given a query protein sequence and NCBI taxonomy IDs (or RefSeq assembly accessions), reactr retrieves genomic data and runs comprehensive analysis across 4 integrated tiers: (1) evolutionary analysis, including homolog detection, domain-based clustering, multiple sequence alignment, and phylogenetic inference; (2) synteny and selection analysis, detecting collinear blocks and calculating Ka/Ks ratios; (3) structural and regulatory characterization, including motif discovery, chromosomal mapping, biochemical property prediction, subcellular localization prediction, and promoter analysis; and (4) experimental design tools, generating PCR primers and scored CRISPR gRNAs for lab validation. Reactr bridges computational prediction and experimental validation, thus enabling rapid transition from genomic discovery to functional studies.

Harmonize numerical values extracted from medical images (e.g. acquired with different models of image-acquisition system)

Browser-based viewer for GenBank and GenPept records — .gb, .gbk, .gbff, .gp and plain GenBank text. Renders an interactive linear and circular feature map, including circular plasmid maps, alongside the annotated source text and the nucleotide/protein sequence. Translates CDS features using the record's own genetic code and translation qualifiers, flags where the stored /translation disagrees with a plain translation, and adds optional computed layers: ORF prediction and restriction-site mapping. Drag a range in the sequence band to select it, then copy that stretch — either strand, as DNA or as protein, plain or FASTA. Filters features by type, handles multi-record files, and keeps open records as local sessions, so a closed tab can be picked up where it was left. Runs entirely in the browser — files are never uploaded.

ArrayAnalysis is a web-based application for transcriptomic data analysis. It supports the analysis of both microarray and RNA-seq data. The tool may also be installed locally as a desktop app, Docker image, or R package.

Supernova is a software package for de novo assembly from Chromium Linked-Reads that are made from a single whole-genome library from an individual DNA source. A key feature of Supernova is that it creates diploid assemblies, thus separately representing maternal and paternal chromosomes over very long distances. Almost all other methods instead merge homologous chromosomes into single incorrect 'consensus' sequences. Supernova is the only practical method for creating diploid assemblies of large genomes.

edf2csv is a local command-line tool for converting EDF, EDF+, BDF, and BDF+ physiological recordings into CSV and JSON files. It exports signal values, channel information, annotations, and recording metadata while preserving original sampling rates, physical units, and discontinuities.

Plotwright is a browser-based statistics and scientific graphing workbench for life-science researchers. It provides a Data to Analysis to Graph workflow for statistical tests, regression and curve fitting, survival and ROC analysis, and editable scientific figures. It can import supported content from modern GraphPad Prism .prism files and reports preserved, approximated, and unsupported objects for review. Projects can be saved as local files. Plotwright is commercial software with a 14-day no-card trial.

Predicting the effect of mutations on protein-RNA binding with Deep Learning | This repository contains all DeepCLIP Python code | A context-aware neural network for modeling and predicting protein binding to nucleic acids using only sequence input | DeepCLIP is a neural network with shallow convolutional layers connected to a bidirectional LSTM layer

Plant Compound Extractor is a desktop application that builds a ready-to-use, deduplicated library of 3D ligand structures for a given plant. It queries multiple natural-product and chemical databases (COCONUT, LOTUS, Wikidata, PubChem, PlantaeDB, USDA Dr. Duke's, KNApSAcK and IMPPAT) in parallel, resolves each compound against PubChem for a canonical structure, and falls back to direct source retrieval when needed. Retrieved structures are then converted to 3D using RDKit, with configurable conformer generation and physicochemical filters (molecular weight, rotatable bonds, ring size, etc.). It can also process a manually supplied compound list, or convert an existing folder of 2D structures to 3D.

Desktop viewer for microscopy and whole slide pathology images on Windows and macOS. Opens whole slide scanner formats (Aperio SVS, Hamamatsu NDPI, MIRAX MRXS, Leica SCN, Ventana BIF) alongside acquisition formats (Zeiss CZI, Nikon ND2, DICOM, TIFF/OME-TIFF) in a single application, providing pyramid navigation of gigapixel images, metadata inspection, calibrated measurements, annotations, and on-device segmentation and object counting.

Pan.bio is a cloud genomics platform for pipeline execution, exploratory analysis, and clinical variant interpretation. Workflows runs validated Nextflow and nf-core pipelines including Sarek, rnaseq, scrnaseq, mag, ampliseq, chipseq and atacseq without local installation. Notebooks provides Python and R sessions with a preinstalled bioinformatics stack, importing public data from GEO, SRA and IPG by accession and reading Workflows outputs directly. VAIC applies ACMG/AMP variant classification with gene-specific rule sets from CanVIG-UK and ClinGen ENIGMA, with automated evidence criteria implemented for BRCA1 and BRCA2. Cohorts provides federated analysis of patient data within a Trusted Research Environment.

Novel machine learning-based method for prediction of ligand binding sites from protein structure.

PhonaLab is a browser-based platform for acoustic analysis of voice recordings, aimed at speech-language pathologists, voice clinicians, and researchers. It computes validated multiparametric acoustic indices — including the Acoustic Voice Quality Index (AVQI), Acoustic Breathiness Index (ABI), smoothed cepstral peak prominence (CPPS), and glottal-to-noise excitation ratio (GNE) — from sustained-vowel and connected-speech recordings, using Praat algorithms via the Parselmouth interface. Audio is processed in memory and not stored. Interface available in English, Brazilian Portuguese, and Spanish.

Gelyze is a web-based image analysis tool for agarose gel electrophoresis of nucleic acids. A user uploads a gel image, confirms the proposed lane positions and the ladder baseline, and a deterministic algorithm then estimates fragment size from ladder calibration and quantifies relative band intensity by densitometry. Measurement and interpretation are kept separate: the numbers come from the algorithm, and an assistive AI layer explains the resulting pattern in plain language without producing any of the values. For research use only.

MailsDaddy MBOX to PST Converter is designed to move single or multiple MBOX file to Outlook PST, EML, MSG, HTML etc.

FT-ITC Analysis is free, open-source desktop software for processing, fitting, and analysing isothermal titration calorimetry (ITC) data. It imports raw data from MicroCal, TA Instruments/NanoAnalyze, and PEAQ ITC project files, as well as integrated heat data and FT-ITC project files. The software supports baseline correction, injection integration, interaction-model fitting, global analysis, uncertainty estimation, advanced thermodynamic analyses, and publication-figure export on macOS, Windows, and Linux.

Browser-based viewer that maps sequencing reads onto one short reference — an amplicon, gene or plasmid. Reads open as Sanger AB1/ABIF, SCF, FASTA, FASTQ or a SAM somebody else already mapped (gzipped files are unpacked in place); the reference as FASTA, GenBank or a read. Both read orientations are tried automatically. The pileup reports per-position depth, where reads disagree with the target, and the consensus — phred-weighted for capillary reads, which keep their chromatogram under the letters. An optional protein lane translates target and consensus side by side. Reads are placed by minimap2 compiled to WebAssembly, or by the built-in aligner. One read or the whole alignment saves as FASTA. Runs entirely in the browser — files are never uploaded.