Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain(1)
Language(1)
License
Source
Type
182 of 7,068 resources
Showing 1–50
shehrozashoaib/LLM_Crystal_CIF
by shehrozashoaibLoRA adapters fine-tuning Qwen2.5-7B-Instruct to emit a full CIF crystal structure from a prompt of reduced composition + target space-group number. Part of a controlled composition-sweep study (MP-20 : MPTS-52 training ratio at fixed volume/steps).
shreyansh12183/olmo2-7b-biomed-chem
by shreyansh12183Vigyan-7B-BioMed-Chem is a domain-specialized LoRA adapter trained on OLMo-2-1124-7B dedicated to organic chemical synthesis, pharmacology, and molecular biology.
helloimsaif/medjev-4b-lora
by helloimsaifTyped-decision adapter for Qwen3.5-4B, tuned on clinical question answering, financial news sentiment and structured record-level workflow decisions.
Taykhoom/Helix-mRNA
by TaykhoomMinimal HuggingFace port of Helix-mRNA -- a hybrid Mamba2 / attention language model for full-length mRNA, trained with next-token prediction on single-nucleotide tokens with a codon-start marker.
insilicomedicine/Qwen3-1.7B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-1.7B, was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-1.7B on aging-related multi-omics and clinical data.
insilicomedicine/longevity-llm
by insilicomedicineA domain-adapted Qwen3.5-9B for aging and longevity biology. L-LLM is the result of continued pretraining + supervised fine-tuning + a reasoning-augmented continuation pass on a multi-domain corpus spanning clinical aging, epigenomics, transcriptomics, proteomics, and genetics.
aigensciences/BioGravity-Inst
by aigensciencesBioGravity-Inst is a biomedical instruction model from AIGEN Sciences, Inc., developed from the Gravity 30B-A5B family. It is intended for biomedical research in the Biomni A1 environment, including question answering, evidence gathering, computation, and tool-assisted analysis.
lighteternal/biodecision-v2-4b
by lighteternalA System-1 decision model for biomedicine, pharma and clinical trials. It reads a source, a question and a set of possible answers, and returns a calibrated probability for each answer in one forward pass, with no generated text.
Soilytix/LOAM-624M
by SoilytixFreedomIntelligence/HuatuoGPT-3-Grader-8B
by FreedomIntelligenceHuatuoGPT-3-Grader-8B GitHub | Paper
Kentucky-Open-Science/KOS-V5-Instruct
by Kentucky-Open-ScienceDeveloped by
Aquiles-ai/Evo2-7B
by Aquiles-aiEvo2-7B (Transformers port)
Aquiles-ai/Evo2-1B-Base
by Aquiles-aiEvo2-1B-Base (Transformers port)
shikunpunk/ask-dao
by shikunpunk> ⚠️ 重要:本仓库的 adapter 历史上因 PeftModel.frompretrained 双重包装导致 key 嵌套错误。 > 旧版本里 PeftModel.frompretrained 加载会"Found missing adapter keys"并静默丢弃全部权重, > 模型实际退化为 base Qwen2.5-3B-Instruct。 > 现在本仓库的 adapter_model.safetensors 已重新打包为标准深度 8(504/504 keys 命中),可被正确加载。 > 验证方式:见 shikunpunk/ask-dao-v0.3 仓库里 "Holdout…
shikunpunk/ask-dao-v0.2-2ep
by shikunpunk知识发现机器 —— 从生物医学论文推断「作者没有明说」的开放科学问题
Aurigene-AI/ChemFM-1B
by Aurigene-AISee the upstream model card for full details, training data and citation.
learning-unit/L1-30B-A5B
by learning-unitL1-30B-A5B is the Korean-locale medical foundation model from Lunit and Lunit Consortium. It is the 30B member of the L1 family, post-trained directly from Gravity-30B-A5B-Base, a sparse Mixture-of-Experts model developed by Trillion Labs and the Lunit Consortium.
CENO-1B-1m is a checkpoint of the CENO base DNA foundation model (1M context (stage 4)). It is a plain causal language model over genomic sequence on a Nemotron-H Mamba/Attention/MoE hybrid backbone, with no MSA inputs.
tzcfly/PertMind
by tzcflyPertMind is a biological language model built around a central discovery: public cellular perturbation atlases can be reorganized into reinforcement-learning environments, where measured gene responses act as computable reward signals for biological reasoning.
NavitraTechnologies01/navikinase-1.0
by NavitraTechnologies01A from-scratch, decoder-only protein language model for the phosphotransferase superfamily (EC 2.7.-: protein kinases plus sugar/lipid/nucleotide kinases), trained entirely locally on Apple Silicon via MLX — no cloud compute, no fine-tuning of an existing model.
prathmeshadsod/BondShift-Llama-3.3-70B-Instruct
by prathmeshadsodBondShift: Organic Mechanism Reasoning
HuggingFaceBio/Carbon-3B
by HuggingFaceBioTechnical Report 🧬
This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…
MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…
insilicomedicine/Qwen3-0.6B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-0.6B, is the smallest family member and was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-0.6B on aging-related multi-omics and…
!Format !Task !Params !Type !License
llmithull/HealthGPT-LoRA
by llmithullHealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.
Meddies/meddies-pii
by MeddiesA multilingual PII extractor for teams that need structured JSON from clinical and administrative text.
ZeroOneAI/ZEO-Med-2
by ZeroOneAITrinity-Mini-AI-Scientist
This is a QLoRA adapter for query-focused structured extraction from one PubMed title and abstract. It was trained as part of BioEvidence Copilot and targets the repository's versioned ModelEvidenceExtraction JSON Schema.
introvoyz041/DrugGen-2
by introvoyz041# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a disease‑aware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
Strict automatic scores on the unchanged 1,309-example primary holdout; compare values within each task panel.
RetroAgent is a 4B-parameter LLM agent for multi-step retrosynthesis planning. It decomposes a target molecule into commercially available building blocks by searching over an AND-OR graph of molecules and reactions, driven entirely by tool calls.
trillionlabs/TxGravity-30B-A5B
by trillionlabsTxGravity-30B-A5B is a therapeutics-focused language model fine-tuned from the Gravity-30B-A5B-base. It is trained to predict a broad range of therapeutic properties — small-molecule ADMET, toxicity, drug–target interaction, protein–protein and peptide–MHC interaction, and more — following the…
alimotahharynia/DrugGen-2
by alimotahharynia# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a disease‑aware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
CladeTeam/CENO-P-1B
by CladeTeamCENO-P-1B is the multi-species alignment (MSA) post-trained variant of the 1B CENO DNA foundation model, for variant effect prediction (VEP). It carries intraencodingpattern in its config and ships the MSA scoring path (modelingcenop.py), which consumes a per-token seq_idx to score packed MSA…
CladeTeam/CENO-1B-131k
by CladeTeamCENO-1B-131k is the long-context (131k) checkpoint of the 1B CENO DNA foundation model — a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
CladeTeam/CENO-80M-1m
by CladeTeamCENO-80M-1m is the long-context (1M) checkpoint of the 80M CENO DNA foundation model — a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B
by reaperdoesntknowA 1.7B-parameter causal language model distilled from Qwen3-30B-A3B on 6,122 STEM chain-of-thought samples using discrepancy-informed knowledge distillation. The training objective emphasizes proof structure, detects reasoning pivot tokens through token-level divergence dynamics, smooths…
fableforge-ai/NEXUS-Medical
by fableforge-ai> NEXUS domain specialist for medical Q&A and clinical reasoning — lightweight & uncensored.
!Screenshot 2026-07-05 at 2.33.47 AM
Pippinlitli/evolva-qwen-0.5b-heretic
by PippinlitliHeretic-abliterated version of Qwen/Qwen2.5-0.5B-Instruct for the Evolva drug discovery pipeline.
zsyjsld/Xinghe1-9B
by zsyjsldXinghe1-9B (杏核) is a specialized large language model fine-tuned for the formalization, computational derivation, and clinical reasoning of Huangdi Neijing. It is based on the Qwen3.5-9B-Instruct architecture and trained using the V3 Double-Purity SFT dataset.