Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type
1,191 of 7,069 resources
Showing 451–500
CliniGuard Clinical Findings NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of clinical findings, diseases, conditions, anatomical locations, and clinical modifiers from unstructured clinical text.
102 executable tasks from 44 peer-reviewed papers across 4 disciplines with containerized evaluation
introvoyz041/DrugGen-2
by introvoyz041# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a disease‑aware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
Reinforcement-learning-trained AI agent for treatment reasoning over a universe of 212 biomedical tools, performing multi-step evidence gathering and spawning parallel reasoning branches to reach evidence-grounded clinical decisions (55+ stars, MIT License, 2026)
Differentiable PDE solving framework for machine learning with built-in fluid simulation, supporting PyTorch/JAX/TensorFlow backends and enabling neural network training within physical simulations (TUM, MIT License)
Strict automatic scores on the unchanged 1,309-example primary holdout; compare values within each task panel.
First scientific ML benchmark with paired real-world measurements and matched numerical simulations for complex physical systems, featuring 5 scenarios, 700+ trajectories, 10 baseline models, and 9 evaluation metrics with HuggingFace datasets and model checkpoints (Westlake University, CC BY-NC 4.0)
The EVORAO Ontology provides a structured and harmonized vocabulary for describing shareable pathogens as characterized biological materials, along with their derived products and associated services, organized into collections. Developed within the EVORA project, it supports consistent metadata annotation across research infrastructures, promoting findability, accessibility, interoperability, and reusability (FAIR). By aligning with relevant standards and ontologies, EVORAO facilitates cross-domain collaboration, integration, and sharing of pathogenic resources and services to enhance pandemic preparedness and response. While initially focused on virology, EVORAO is designed to be extensible and also supports metadata harmonization for other pathogens. [from repository]
RetroAgent is a 4B-parameter LLM agent for multi-step retrosynthesis planning. It decomposes a target molecule into commercially available building blocks by searching over an AND-OR graph of molecules and reactions, driven entirely by tool calls.
RFdiffusion is an open source method for structure generation, with or without conditional information (a motif, target etc).
trillionlabs/TxGravity-30B-A5B
by trillionlabsTxGravity-30B-A5B is a therapeutics-focused language model fine-tuned from the Gravity-30B-A5B-base. It is trained to predict a broad range of therapeutic properties — small-molecule ADMET, toxicity, drug–target interaction, protein–protein and peptide–MHC interaction, and more — following the…
PathBench-MIL is a comprehensive, flexible benchmarking/AutoML framework for multiple instance learning in histopathology. PathBench-MIL is expected to be deprecated and replaced by PathForge.
Py-HLA-Match is a Python library for standardised, rule-based HLA (Human Leukocyte Antigen) matching in retrospective analyses, method development, benchmarking, and in-silico studies in immunogenetics and related fields.
AI-powered field boundary delineation toolkit combining satellite foundation models, embeddings, and global training data for accurate agricultural parcel/field boundary mapping, with Google Earth Engine integration and PyPI distribution (84+ stars, Apache 2.0, 2026)
Autonomous AI agent for end-to-end spatial proteomics analysis, featuring SP-Bench for agentic multiplexed-imaging workflows (tomtommyyuan, 140+ stars, 2026)
Phsntom/ESMFold2-Fast
by PhsntomESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…
alimotahharynia/DrugGen-2
by alimotahharynia# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a disease‑aware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
CladeTeam/CENO-P-1B
by CladeTeamCENO-P-1B is the multi-species alignment (MSA) post-trained variant of the 1B CENO DNA foundation model, for variant effect prediction (VEP). It carries intraencodingpattern in its config and ships the MSA scoring path (modelingcenop.py), which consumes a per-token seq_idx to score packed MSA…
CladeTeam/CENO-1B-131k
by CladeTeamCENO-1B-131k is the long-context (131k) checkpoint of the 1B CENO DNA foundation model — a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
CladeTeam/CENO-80M-1m
by CladeTeamCENO-80M-1m is the long-context (1M) checkpoint of the 80M CENO DNA foundation model — a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
Universal Cell Embeddings: zero-shot single-cell foundation model pretrained on 36M cells across 11M species, learning cross-species gene function representations via a protein-language-model-informed token space; enables zero-shot cell type annotation, embedding, and integration of unseen datasets and species without fine-tuning (338+ stars, MIT License)
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B
by reaperdoesntknowA 1.7B-parameter causal language model distilled from Qwen3-30B-A3B on 6,122 STEM chain-of-thought samples using discrepancy-informed knowledge distillation. The training objective emphasizes proof structure, detects reasoning pivot tokens through token-level divergence dynamics, smooths…
DeepTaxa is a hybrid CNN-BERT deep learning framework for multi-rank taxonomic classification of 16S rRNA gene sequences. It predicts all seven Linnaean ranks from domain to species in a single forward pass and provides pre-trained checkpoints for full-length 16S and V3-V4 amplicons.
Fine-tuned ESM-2 650M with LoRA for predicting protein subcellular localization (10 classes).
phenobase/phenovisionL
by phenobasePhenoVisionL is a Vision Transformer (ViT-Large) model fine-tuned to detect leaf phenological states in plant photographs: green leaves, colored (senescent) leaves, and breaking leaf buds. It was trained on 165,988 iNaturalist records of deciduous woody plants using a two-stage semi-supervised…
phenobase/phenovision
by phenobasePhenoVision is a Vision Transformer (ViT-Large) model fine-tuned to detect flowers and fruits in plant photographs. It was trained on 1.5 million human-annotated iNaturalist images and has been used to generate over 30 million new phenology records across 119,000+ plant species, vastly expanding…
fableforge-ai/NEXUS-Medical
by fableforge-ai> NEXUS domain specialist for medical Q&A and clinical reasoning — lightweight & uncensored.
Clinical-Reasoning-Hub/pentabrid-27b
by Clinical-Reasoning-HubA package for creating fast and accurate interatomic potentials.
lowdown-labs/fela-genomics
by lowdown-labslowdown-labs/fela-chemistry
by lowdown-labsAn interactive platform that performs statistical analyses on metabolomics datasets and allows visualising results with ease. The interface gives users autonomy in creating figures suited to their reporting and publication needs.
Long-context generative genomic foundation model using 6-mer tokenization for DNA sequence modeling and generation, with v2 model families for prokaryote and eukaryote genomes and pretrained weights available on HuggingFace (GenerTeam, 460+ stars, MIT License, 2025-2026)
!Screenshot 2026-07-05 at 2.33.47 AM
This repository contains LoRA finetunes of DiffusionGemma (image-conditioned discrete-diffusion LLM) for radiology visual question answering, each paired with an autoregressive Gemma-4 finetune as a controlled baseline. It corresponds to the paper Discrete Diffusion Language Models for Interactive…
Multi-agent system for drug-discovery gene target validation. LangGraph agents over an MCP data layer (~26 data sources, ~44 tools) score evidence across six independent lenses (genetics, biology, safety, clinical, commercial, regulatory) into a provenanced dossier. Configurable local/cloud LLM routing with full Langfuse/OTEL traceability.
Democratizing AlphaFold3: PyTorch reimplementation to accelerate protein structure prediction research
doctolib-lab/doctobert-fr-base
by doctolib-lab🤗 Blog | 📄 Paper | 💻 Code | 🌐 FineMed | 🩺 DoctoBERT
Medical time series foundation model pretrained on 454B time points from heterogeneous clinical corpora spanning ICU physiological signals and hospital EHR, with continuous-time rotary positional encoding, frequency-specialized Mixture-of-Experts, and neural ODE extrapolation for zero-shot forecasting across irregular and multimodal temporal health data (Microsoft, 399+ stars, MIT License)
Parameter/topology editor and molecular simulator with visualization capability.
Pippinlitli/evolva-qwen-0.5b-heretic
by PippinlitliHeretic-abliterated version of Qwen/Qwen2.5-0.5B-Instruct for the Evolva drug discovery pipeline.
HantaBERT/HantaBERT
by HantaBERTHantaBERT fine-tunes DNABERT-2 on hantavirus RNA sequences for three simultaneous classification tasks: species/lineage, host, and geographic origin. A single forward pass produces predictions for all three tasks along with a 768-dimensional embedding suitable for phylogenetic visualization.
mradermacher/CellHermes-v1.0-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
All-atom generative world model for all-to-all biomolecular interaction design, enabling cross-modality generation of proteins, nucleic acids, small molecules, and cyclic peptides with fine-grained epitope-level control and 2-4 orders of magnitude faster design throughput than modality-specific baselines (316+ stars, Apache 2.0)
zsyjsld/Xinghe1-9B
by zsyjsldXinghe1-9B (杏核) is a specialized large language model fine-tuned for the formalization, computational derivation, and clinical reasoning of Huangdi Neijing. It is based on the Qwen3.5-9B-Instruct architecture and trained using the V3 Double-Purity SFT dataset.