Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type
1,191 of 7,068 resources
Showing 151–200
biohub/ESMFold2
by biohubESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…
biohub/ESMFold2-Fast
by biohubESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…
scrc-dnai/DNT-8M-CPL-SMP
by scrc-dnaiscrc-dnai/DNT-100M_PRE-CPL-2e4
by scrc-dnaiOffline tool for cleaning tables of human gene and protein identifiers (TXT, CSV, TSV, XLSX). It maps approved symbols, aliases, previous symbols, Ensembl gene, UniProt, Entrez, RefSeq and HGNC identifiers to current HGNC approved symbols with cross-references, using a bundled HGNC snapshot. Excel date-corrupted symbols are recovered where the original is unambiguous and flagged for manual review otherwise; no input row is dropped. Each run records the tool version and HGNC release and writes a per-row audit table.
Aquiles-ai/Evo2-7B
by Aquiles-aiEvo2-7B (Transformers port)
Aquiles-ai/Evo2-1B-Base
by Aquiles-aiEvo2-1B-Base (Transformers port)
Google DeepMind's official collection of agentic science skills accelerating scientific workflows with better grounding and higher token efficiency, integrating insights from AlphaGenome, AFDB, UniProt and 30+ other databases and tools (2026)
Universal pretrained neural network potential with charge and magnetic moment awareness, trained on 1.5M+ Materials Project inorganic structures for charge-informed molecular dynamics and phase diagram prediction (Berkeley, Nature Machine Intelligence 2023 Cover)
Python framework for writing high-performance GPU simulation and graphics kernels with first-class automatic differentiation, enabling differentiable physics, molecular dynamics, soft-body and cloth simulation, robotics, and CFD adjoints compiled to CUDA; serves as the underlying engine for differentiable simulation projects like Newton and integrates with PyTorch, JAX, and OpenUSD (NVIDIA, 7.1K+ stars, Apache 2.0)
PyTorch-based differentiable programming framework for physics-informed system identification, parametric constrained optimization, and model predictive control, integrating neural operators, neural ODEs, KANs, SINDy, and differentiable predictive control with 30+ tutorials (1.3k+ stars, BSD License)
An ontology with predicates to formalization of the concept of mentions. The mentions may be either explicit (e.g. as when well stated into an article "Dr. Johnson's groundbreaking research on climate change") or implicit (e.g. such as discussing "seminal studies in the field"). MiTO contains the object property mito:mentions and its inverse mito:isMentionedBy.
Microsoft's foundation model for the Earth system supporting weather, air pollution, and ocean wave forecasting at multiple resolutions, trained on 1M+ hours of diverse atmospheric data (Nature 2025)
Minimalist, batteries-included repository for training video world models with diffusion-forcing, supporting long-horizon rollouts, 3D point-cloud generation, and model-predictive control with pretrained checkpoints (Simchowitz Lab, 700+ stars, MIT License, 2026)
An EMMO-based domain ontology for atomistic and electronic modelling.
An ontology that provides a structured vocabulary for rhetorical elements within documents (e.g., Introduction, Discussion, Acknowledgements, Reference List, Figures, Appendix). It is imported by DoCO.
Generates pre-miRNA and mature miRNA count tables from read alignments to pre-miRNA sequences and a gff file, both downloaded from mirBase. Produces also read coverage plots of pre-miRNAs.
PyTorch-native atomistic simulation engine for the machine-learned interatomic potential (MLIP) era, enabling batched molecular dynamics and structural relaxation with automatic GPU memory management; supports MACE, Fairchem, SevenNet, ORB, MatterSim and other popular MLIPs with up to 100x speedup over ASE (Radical AI, AI for Science 2026, 468+ stars, MIT License)
Modular framework for AI-driven scientific and algorithmic discovery, providing a unified interface for implementing, running, and fairly comparing discovery algorithms across 200+ optimization tasks; introduces AdaEvolve and EvoX adaptive/evolutionary algorithms and natively supports OpenEvolve, GEPA, and Harbor-format benchmarks (skydiscover-ai, 568+ stars, Apache 2.0, 2026)
Open source PEM (Proton Exchange Membrane) fuel cell simulation tool.
Open-source SDK for working with quantum computers at the level of extended quantum circuits, operators, and primitives, enabling quantum algorithm development for quantum chemistry, materials science, and optimization research (IBM, 7.4K+ stars, Apache 2.0)
An ontology that allows the description of numerical and categorical bibliometric data (e.g., journal impact factor, author h-index, categories describing research careers) in RDF.
High-accuracy PDF→Markdown/JSON/HTML conversion, specialized for tables/formulas/code blocks with benchmark scripts
HIDE-Deconv is a framework for characterizing cellular remodeling from bulk RNA-seq data using hierarchical cell-type deconvolution across multiple levels of cellular resolution. It provides an integrated workflow for single-cell reference preprocessing, estimation of cellular compositions, and downstream analysis of deconvolution results, including visualization, clustering, differential composition, and survival analysis.
Evolvable and privacy-preserving multi-agent framework automating, scaling, and accelerating data sciences with a particular focus on end-to-end single-cell biology analyses; features agentic code evolution, multi-agent team orchestration, distributed architecture, and a community marketplace with 1,000+ curated agents and skills (428+ stars)
ctheodoris/Geneformer
by ctheodoris# Geneformer Geneformer is a foundational transformer model pretrained on a large-scale corpus of human single cell transcriptomes to enable context-aware predictions in settings with limited data in network biology.
Terms for genes, experimental factors, and cell lines used by the [Gemma platform](https://gemma.msl.ubc.ca/home.html) for differential gene expression analysis.
Flow-matching protein folding model using only general-purpose transformer layers, scaled to 3B parameters and trained on 8.6M+ distilled structures; challenges the reliance on complex domain-specific architectures and supports PyTorch and MLX backends with model sizes from 100M to 3B parameters (985+ stars, MIT License)
Large-scale knowledge graph and pip-installable client for literature-grounded automated scientific research, connecting papers, authors, institutions, venues, keywords, citations, and a four-level research taxonomy across medicine, social sciences, engineering, computer science, materials science, and more (ZJU NLP, arXiv 2026, 136+ stars, MIT License)
The Ontology of Immune Epitopes (ONTIE) is an effort to represent terms in the immunology domain in a formal ontology with the specific goal of representing experiments that identify and characterize immune epitopes.
Ontology representation of the [International Committee on Taxonomy of Viruses (ICTV)](https://ictv.global/) for the [EVORA project](https://evora-project.eu/)
Deterministic, rule-based variant interpretation platform for clinical genetics laboratories. Automates ACMG/AMP 2015 classification using a Bayesian point-based framework (Tavtigian et al. 2018) with BayesDel ClinGen SVI-calibrated thresholds (Pejaver et al. 2022). Integrates 8 reference databases (gnomAD v4.1, ClinVar, dbNSFP 4.9c, SpliceAI, gnomAD Constraint, HPO, ClinGen, Ensembl VEP). Analyzes nuclear and mtDNA variants, structural and copy-number variants (SV/CNV), with trio/family and cohort analysis. Supports HPO-based phenotype matching, biomedical literature mining across 2M+ PubMed publications, and structured clinical report generation. AI assists in evidence synthesis but does not make classification decisions. EU-hosted on dedicated infrastructure in Helsinki, Finland (GDPR-compliant).
Fully open-source (Apache 2.0) biomolecular structure prediction reproducing AlphaFold3, free for academic and commercial use (Columbia AlQuraishi Lab & OpenFold Consortium, 2025)
RBPBench is a multi-function tool to evaluate CLIP-seq and other related genomic region data using a comprehensive collection of known RNA-binding protein (RBP) binding motifs. RBPBench can be used for a variety of purposes, from RBP motif search (database or user-supplied RBP motifs) in genomic regions, over motif enrichment and co-occurrence analysis, in-depth comparisons over multiple datasets via sequence and genomic annotation statistics, to benchmarking CLIP-seq peak caller methods as well as comparisons across cell types and CLIP-seq protocols. RBPBench supports both sequence and structure motifs, as well as regular expressions (sequence and structure patterns). Moreover, users can easily provide their own motif collections.
Microsoft's AI-powered geospatial Earth science application for natural-language exploration, visualization, and analysis of 130+ satellite collections, with STAC integration, multi-agent backend, MCP server, and deployable React/FastAPI stack (MIT, 2025)
Deep probabilistic framework for single-cell and spatial omics analysis, integrating scVI, scANVI, totalVI and other VAE-based models for batch correction, cell annotation, multi-omics integration, and RNA velocity (scverse/NumFOCUS, Nature Methods 2018/2024)
The "FRamewOrk for Molecular AGgregate Excitations" enables localised QM/QM' excited state calculations in a solid state environment.
shikunpunk/ask-dao
by shikunpunk> ⚠️ 重要:本仓库的 adapter 历史上因 PeftModel.frompretrained 双重包装导致 key 嵌套错误。 > 旧版本里 PeftModel.frompretrained 加载会"Found missing adapter keys"并静默丢弃全部权重, > 模型实际退化为 base Qwen2.5-3B-Instruct。 > 现在本仓库的 adapter_model.safetensors 已重新打包为标准深度 8(504/504 keys 命中),可被正确加载。 > 验证方式:见 shikunpunk/ask-dao-v0.3 仓库里 "Holdout…
shikunpunk/ask-dao-v0.2-2ep
by shikunpunk知识发现机器 —— 从生物医学论文推断「作者没有明说」的开放科学问题
Aurigene-AI/ChemFM-1B
by Aurigene-AISee the upstream model card for full details, training data and citation.
learning-unit/L1-30B-A5B
by learning-unitL1-30B-A5B is the Korean-locale medical foundation model from Lunit and Lunit Consortium. It is the 30B member of the L1 family, post-trained directly from Gravity-30B-A5B-Base, a sparse Mixture-of-Experts model developed by Trillion Labs and the Lunit Consortium.
JAMMA (Highly-Accelerated Multi-method Mixed-Model Association) is an open-source Python and C implementation of GEMMA's core linear mixed-model workflows for genome-wide association studies (GWAS). It reads PLINK binary genotypes and supports kinship estimation, Wald, likelihood-ratio and score association tests, covariates, multiple phenotypes, and leave-one-chromosome-out (LOCO) analysis. JAMMA provides a command-line interface with familiar GEMMA flags, GEMMA-compatible association output, and a Python API. Native C kernels, parallel computation and reusable eigendecompositions support large analyses. Pre-flight memory checks and streamed output help manage memory use. Numerical validation against GEMMA is documented. JAMMA runs on Linux, macOS and Windows; large-cohort analyses require sufficient RAM and a suitable 64-bit BLAS configuration. Released under GPL-3.0-or-later.
ChemML is a machine learning and informatics program suite for the analysis, mining, and modeling of chemical and materials data. (based on Tensorflow)
The Open Forcefield Toolkit provides implementations of the SMIRNOFF format, parameterization engine, and other tools.
A batteries-included toolkit for the GPU-accelerated OpenMM molecular simulation engine.
Google Research's hybrid ML/physics atmospheric model combining learned dynamics with physical constraints, outperforming traditional models on 2-15 day forecasts and 40-year climate simulation, developed with ECMWF (Nature 2024)
- Molecular Manipulation Made Easy. A light wrapper build on top of RDKit.
pyLocusZoom is an open-source Python library for visualizing genome-wide association study (GWAS) results. It creates LocusZoom-style regional association plots with linkage disequilibrium (LD) coloring, gene and exon tracks, and recombination overlays. Additional plots include Manhattan, QQ, Miami, eQTL, fine-mapping credible sets, PheWAS, forest plots, LD heatmaps and colocalization comparisons. Matplotlib provides static figures; Plotly and Bokeh provide interactive views. The library accepts pandas DataFrames and includes loaders for PLINK, GEMMA, REGENIE, BOLT-LMM, SAIGE, GTEx, SuSiE and FINEMAP outputs. It supports canine and feline reference data, automatic Ensembl gene annotations, and custom reference data for other species. LD can be supplied or calculated with PLINK. Requires Python 3.10 or later and is released under GPL-3.0-or-later.
The Essential FRBR in OWL2 DL Ontology (FRBR) is an expression in OWL 2 DL of the basic concepts and relations described in the IFLA report on the Functional Requirements for Bibliographic Records (FRBR), also described in Ian Davis's RDF vocabulary. It is imported by FaBiO and BiRO.
UniParser/MolParser-Mobile
by UniParser💻 Github | 📄 Report | 🚀 Demo