Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type
1,191 of 7,068 resources
Showing 1–50
shehrozashoaib/LLM_Crystal_CIF
by shehrozashoaibLoRA adapters fine-tuning Qwen2.5-7B-Instruct to emit a full CIF crystal structure from a prompt of reduced composition + target space-group number. Part of a controlled composition-sweep study (MP-20 : MPTS-52 training ratio at fixed volume/steps).
GPU-accelerated differentiable physics simulation engine built on NVIDIA Warp, supporting rigid/soft body, cloth, and gradient-based optimization for scientific ML, initiated by Disney Research, DeepMind, and NVIDIA (Linux Foundation, Apache 2.0, 2025)
HuggingFaceBio/Carbon-A-1.2B
by HuggingFaceBioA DNA annotation model from the Carbon family.
pantoniadis/RAGenome
by pantoniadisPaper: RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts
Open-source framework for building physics-ML models at scale (renamed from Modulus, 2025)
PyTorch library for training neural networks on gravitational-wave physics, providing differentiable PSD estimation, whitening, SNR calculation, interferometer response projection, waveform simulation, and streaming data loaders; the shared back end of the NSF A3D3/ML4GW pipelines deployed for real-time detection of compact-binary coalescences in LIGO–Virgo–KAGRA observing runs (35+ stars, GPL-3.0, 2022-2026)
MedDecider-27B is the most robust mid-size member of the MedDecider family. Given a clinical state (a note, a trial record or any JSON) and a question with 2 to 10 answer options, it returns a calibrated probability for every option in a single scoring pass, without generating text.
Automated pipeline for proteome-scale protein-protein interaction screening with AlphaFold-Multimer and AlphaFold 3, supporting flexible inputs (UniProt IDs, FASTA, residue regions, multimers, AF3 JSON features) and integrated downstream analysis for hit prioritization (Kosinski Lab, EMBL, Nature Protocols 2024, 317+ stars, GPL-3.0)
shreyansh12183/olmo2-7b-biomed-chem
by shreyansh12183Vigyan-7B-BioMed-Chem is a domain-specialized LoRA adapter trained on OLMo-2-1124-7B dedicated to organic chemical synthesis, pharmacology, and molecular biology.
helloimsaif/medjev-4b-lora
by helloimsaifTyped-decision adapter for Qwen3.5-4B, tuned on clinical question answering, financial news sentiment and structured record-level workflow decisions.
Open software framework for Engineering AI built on transformer building blocks, enabling teams to build, train, and operate industrial simulation models across engineering verticals; includes ready-to-use recipes for CFD (AB-UPT on DrivAerML), external aerodynamics, and heat transfer (234+ stars, ENPL non-commercial license, 2026)
Research ecosystem for rigorous and trustworthy AI scientists — a protocol and skill bundle that makes autonomous research verifiable, crystallized, and observable through structured, machine-executable research artifacts and five agent skills for research management, compilation, verification, visualization, and publication (ARA-Labs, 447+ stars, MIT License, 2026)
Cross-platform library for differentiable programming of quantum computers with automatic differentiation, enabling hybrid quantum-classical machine learning for quantum chemistry, quantum physics, and NISQ algorithm research (Xanadu, 3k+ stars)
Taykhoom/Helix-mRNA
by TaykhoomMinimal HuggingFace port of Helix-mRNA -- a hybrid Mamba2 / attention language model for full-length mRNA, trained with next-token prediction on single-nucleotide tokens with a codon-start marker.
spoQC is a modular framework for multimodal quality control (QC) of imaging-based spatially resolved transcriptomics (SRT). It independently evaluates cell segmentation, imaging, and transcript data to identify high-quality regions (HQRs) across entire tissue sections. In addition, spoQC uses Markov random fields (MRFs) to incorporate spatial dependencies and generate spatially refined QC masks.
Composite-objective protein design framework integrating Boltz, AlphaFold2, OpenFold3, ProteinMPNN, and ESM via JAX-based gradient optimization over continuous relaxed sequence space for multi-property binder design (319+ stars, MIT License, 2025)
Unified framework for state-of-the-art pre-trained bio foundation models across genomics and transcriptomics, providing standardized interfaces and pipelines for DNA, RNA, and single-cell models including Evo 2, Geneformer, scGPT, and UCE with streamlined inference, benchmarking, and fine-tuning workflows (213+ stars, 2024-2025)
insilicomedicine/Qwen3-1.7B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-1.7B, was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-1.7B on aging-related multi-omics and clinical data.
insilicomedicine/longevity-llm
by insilicomedicineA domain-adapted Qwen3.5-9B for aging and longevity biology. L-LLM is the result of continued pretraining + supervised fine-tuning + a reasoning-augmented continuation pass on a multi-domain corpus spanning clinical aging, epigenomics, transcriptomics, proteomics, and genetics.
Transformer encoder-decoder for de novo peptide sequencing from tandem mass spectrometry, translating MS/MS spectra directly to peptide sequences without reference databases, enabling identification of novel peptides for immunopeptidomics, antibody repertoires, and metaproteomes (Noble Lab UW, Nature Communications 2024)
Interactive and hardware-agnostic SDK for laboratory automation, enabling programmatic control of liquid handlers, plate readers, and other lab instruments across multiple vendors; foundational infrastructure for self-driving laboratories and AI-driven experimental execution (447+ stars)
aigensciences/BioGravity-Inst
by aigensciencesBioGravity-Inst is a biomedical instruction model from AIGEN Sciences, Inc., developed from the Gravity 30B-A5B family. It is intended for biomedical research in the Biomni A1 environment, including question answering, evidence gathering, computation, and tool-assisted analysis.
Open-source Bayesian optimization and design-of-experiments framework serving as the optimization back end of self-driving laboratory campaigns, including the AlphaFlow autonomous synthesis platform (Nature 2024); provides surrogate models, active/transfer learning strategies, chemistry-aware encodings (RDKit fingerprints, descriptors), and botorch-based uncertainty handling with a unified, pip-installable API (513+ stars, Apache 2.0, 2023-2026)
lighteternal/biodecision-v2-4b
by lighteternalA System-1 decision model for biomedicine, pharma and clinical trials. It reads a source, a question and a set of possible answers, and returns a calibrated probability for each answer in one forward pass, with no generated text.
Multimodal deep learning framework integrating peptide-MHC protein sequence, structure, and biochemical properties to predict class-I immunogenicity for infectious disease epitopes and cancer neoepitopes with cancer-wildtype contrastive learning, enabling personalized vaccine design (Krishnaswamy Lab, Yale University)
PathForge is a modular benchmarking framework for multiple instance learning in computational pathology. It supports whole slide image feature extraction, HDF5 artifact generation, tile overviews, benchmarking, pipeline optimization, classification, regression, survival and retrieval tasks, and support for model inference and visualization.
AI coding agent skills for KiCad electronics design that turn Claude Code, Codex, Gemini CLI, and other coding agents into full electronics design assistants; parses schematics and PCB layouts, builds power trees, audits connectors/ESD protection, validates passive networks, runs SPICE simulation, sources components from major distributors, and prepares boards for fabrication (aklofas, 974+ stars, MIT License, 2026)
athanzli/MicroGlot
by athanzliA taxonomy-informed sparse DNA foundation model for microbial genomics.
A Python package for protein dynamics analysis
A data model for managing information about chemical entities, ranging from atoms through molecules to complex mixtures.
drzo/ESMC-6B
by drzoESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…
Hand-curated Snakemake pipelines to combine identifier cross-references from multiple sources across dozens of biomedical types, including anatomical entities, diseases and phenotypes, genes and proteins and many others.
Co-create PowerPoint presentations with Generative AI from documents or topics
Open-source image analysis toolkit for high-throughput plant phenotyping, extracting morphological, color, and texture traits from RGB, hyperspectral, and thermal imagery with modular Python workflows for crop improvement, stress detection, and plant biology research (Donald Danforth Plant Science Center, 795+ stars, MPL-2.0)
Molecular dynamics analysis
Ensemble of automated machine learning protocols that can be run sequentially through a single command line. The program works for regression and classification problems.
An ontology in the OBO foundry, not exactly the same as the obo namespace
The Common Core Ontologies (CCO) comprise twelve ontologies that are designed to represent and integrate taxonomies of generic classes and relations across all domains of interest. CCO is a mid-level extension of Basic Formal Ontology (BFO), an upper-level ontology framework widely used to structure and integrate ontologies in the biomedical domain (Arp, et al., 2015). BFO aims to represent the most generic categories of entity and the most generic types of relations that hold between them, by defining a small number of classes and relations. CCO then extends from BFO in the sense that every class in CCO is asserted to be a subclass of some class in BFO, and that CCO adopts the generic relations defined in BFO (e.g., has_part) (Smith and Grenon, 2004). Accordingly, CCO classes and relations are heavily constrained by the BFO framework, from which it inherits much of its basic semantic relationships.
First bioinformatics-native AI agent skill library enabling local-first, reproducible genomic and population-genetics research workflows built on OpenClaw (871+ stars, MIT License, 2026)
PyTorch framework for training neural network interatomic potentials with the Equivariant Transformer (ET) architecture and its efficient TensorNet successor, providing equivariant message passing with linear complexity in tensor order; underpins the MACE-OFF and SPICE models and widely adopted across molecular dynamics and materials simulation workflows (Amsterdam Machine Learning Lab / De Fabritiis Group, 483+ stars, MIT License, actively maintained)
For a convenient overview and download list, visit our model page for this model.
For a convenient overview and download list, visit our model page for this model.
The Context and Measurement Ontology (COMO) contains ontological terms to describe the context for various types of experimental data and measurements. It is useful in its current state for several different environmental microbiology projects. This ontology is used in multiple CORAL (Contextual Ontology-based Repository Analysis Library) deployments.
Descriptor computation(chemistry) and (optional) storage for machine learning.
Continuously updated functional re-annotation of the Mycobacterium tuberculosis complex gene set, anchored on the MTBC0 ancestral genome rather than on a single strain. Serves one record per gene combining Pfam domains, ESMFold structures with Foldseek search, protein language-model features, orthology, curated knowledge, protein association networks and intra-species selection inferred from 145209 sequenced genomes, with dated sources and a graded confidence level for every field. Intended as a successor to Mycobrowser, which is no longer maintained.