Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain(1)
Language
License
Source
Type
228 of 7,050 resources
Showing 1–50
Taykhoom/Helix-mRNA
by TaykhoomMinimal HuggingFace port of Helix-mRNA -- a hybrid Mamba2 / attention language model for full-length mRNA, trained with next-token prediction on single-nucleotide tokens with a codon-start marker.
insilicomedicine/longevity-llm
by insilicomedicineA domain-adapted Qwen3.5-9B for aging and longevity biology. L-LLM is the result of continued pretraining + supervised fine-tuning + a reasoning-augmented continuation pass on a multi-domain corpus spanning clinical aging, epigenomics, transcriptomics, proteomics, and genetics.
aigensciences/BioGravity-Inst
by aigensciencesBioGravity-Inst is a biomedical instruction model from AIGEN Sciences, Inc., developed from the Gravity 30B-A5B family. It is intended for biomedical research in the Biomni A1 environment, including question answering, evidence gathering, computation, and tool-assisted analysis.
lighteternal/biodecision-v2-4b
by lighteternalA System-1 decision model for biomedicine, pharma and clinical trials. It reads a source, a question and a set of possible answers, and returns a calibrated probability for each answer in one forward pass, with no generated text.
joelinator/dflow-novo-model
by joelinatorDFlowNovo is a state-of-the-art de novo peptide sequencing system powered by Continuous-Time Markov Chain Discrete Flow Matching (CTMC-DFM). By treating peptide sequencing as continuous probability flows over discrete amino acid states and integrating dynamic programming (KnapsackDP) reachability…
Soilytix/LOAM-624M
by SoilytixFreedomIntelligence/HuatuoGPT-3-Grader-8B
by FreedomIntelligenceHuatuoGPT-3-Grader-8B GitHub | Paper
Kentucky-Open-Science/KOS-V5-Instruct
by Kentucky-Open-ScienceDeveloped by
endless-frontier/Fx-Bio
by endless-frontierFx-Bio-0913 is a biomedical reasoning large language model post-trained on DeepSeek-V4-Flash, developed by The Endless Frontier lab. It is specialized for biological and biomedical research tasks — including gene-function puzzles, experimental reasoning, and multi-step evidence integration —…
shreyansh12183/olmo2-7b-biomed-chem
by shreyansh12183A domain-adapted small language model specialized for molecular conformation analysis, computational biochemistry, and pharmacology reasoning.
Aquiles-ai/Evo2-7B
by Aquiles-aiEvo2-7B (Transformers port)
Aquiles-ai/Evo2-1B-Base
by Aquiles-aiEvo2-1B-Base (Transformers port)
!Benchmark card: Hertz 0.7F vs same-size models
shikunpunk/ask-dao
by shikunpunk> ⚠️ 重要:本仓库的 adapter 历史上因 PeftModel.frompretrained 双重包装导致 key 嵌套错误。 > 旧版本里 PeftModel.frompretrained 加载会"Found missing adapter keys"并静默丢弃全部权重, > 模型实际退化为 base Qwen2.5-3B-Instruct。 > 现在本仓库的 adapter_model.safetensors 已重新打包为标准深度 8(504/504 keys 命中),可被正确加载。 > 验证方式:见 shikunpunk/ask-dao-v0.3 仓库里 "Holdout…
shikunpunk/ask-dao-v0.2-2ep
by shikunpunk知识发现机器 —— 从生物医学论文推断「作者没有明说」的开放科学问题
Aurigene-AI/ChemFM-1B
by Aurigene-AISee the upstream model card for full details, training data and citation.
learning-unit/L1-30B-A5B
by learning-unitL1-30B-A5B is the Korean-locale medical foundation model from Lunit and Lunit Consortium. It is the 30B member of the L1 family, post-trained directly from Gravity-30B-A5B-Base, a sparse Mixture-of-Experts model developed by Trillion Labs and the Lunit Consortium.
InstaDeepAI/instanovo-phospho-v1.0.0
by InstaDeepAIInstaNovo-P is a specialized transformer-based model for de novo peptide sequencing from phosphoproteomics mass spectrometry data. This model is specifically trained and optimized for identifying phosphorylated peptides and their modification sites.
litert-community/MedGemma-1.5-4B-IT
by litert-communityCENO-1B-1m is a checkpoint of the CENO base DNA foundation model (1M context (stage 4)). It is a plain causal language model over genomic sequence on a Nemotron-H Mamba/Attention/MoE hybrid backbone, with no MSA inputs.
tzcfly/PertMind
by tzcflyPertMind is a biological language model built around a central discovery: public cellular perturbation atlases can be reorganized into reinforcement-learning environments, where measured gene responses act as computable reward signals for biological reasoning.
NavitraTechnologies01/navikinase-1.0
by NavitraTechnologies01A from-scratch, decoder-only protein language model for the phosphotransferase superfamily (EC 2.7.-: protein kinases plus sugar/lipid/nucleotide kinases), trained entirely locally on Apple Silicon via MLX — no cloud compute, no fine-tuning of an existing model.
Run locally - Benchmarks - Whitepaper - Model details - Responsible use
prathmeshadsod/BondShift-Llama-3.3-70B-Instruct
by prathmeshadsodBondShift: Organic Mechanism Reasoning
HuggingFaceBio/Carbon-3B
by HuggingFaceBioTechnical Report 🧬
This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…
MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…
insilicomedicine/Qwen3-1.7B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-1.7B, was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-1.7B on aging-related multi-omics and clinical data.
insilicomedicine/Qwen3-0.6B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-0.6B, is the smallest family member and was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-0.6B on aging-related multi-omics and…
!Format !Task !Params !Type !License
!Format !Task !Params !Type !License
lemuralabs/Gemma-2B-Uncensored
by lemuralabs!Format !Task !Params !Type !License
llmithull/HealthGPT-LoRA
by llmithullHealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.
Meddies/meddies-pii
by MeddiesA multilingual PII extractor for teams that need structured JSON from clinical and administrative text.
ZeroOneAI/ZEO-Med-2
by ZeroOneAIHaamipromax/HamAI-Science-1b
by HaamipromaxA lightweight language model designed to answer science questions clearly and accurately in English.
Trinity-Mini-AI-Scientist
Billy-Liu-DUT/OmniChem-7B-v1
by Billy-Liu-DUTOmniChem is a new series of large language models specialized for the domain of chemistry. It is designed to address the critical challenge of model hallucination in scientific applications. For OmniChem, we release this 7B instruction-tuned model with strong reasoning capabilities.
This is a QLoRA adapter for query-focused structured extraction from one PubMed title and abstract. It was trained as part of BioEvidence Copilot and targets the repository's versioned ModelEvidenceExtraction JSON Schema.
introvoyz041/DrugGen-2
by introvoyz041# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a disease‑aware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
Strict automatic scores on the unchanged 1,309-example primary holdout; compare values within each task panel.
RetroAgent is a 4B-parameter LLM agent for multi-step retrosynthesis planning. It decomposes a target molecule into commercially available building blocks by searching over an AND-OR graph of molecules and reactions, driven entirely by tool calls.
trillionlabs/TxGravity-30B-A5B
by trillionlabsTxGravity-30B-A5B is a therapeutics-focused language model fine-tuned from the Gravity-30B-A5B-base. It is trained to predict a broad range of therapeutic properties — small-molecule ADMET, toxicity, drug–target interaction, protein–protein and peptide–MHC interaction, and more — following the…