Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain(1)
Language(1)
License
Source
Type
149 of 6,573 resources
Showing 1β50
tzcfly/PertMind
by tzcflyPertMind is a biological language model built around a central discovery: public cellular perturbation atlases can be reorganized into reinforcement-learning environments, where measured gene responses act as computable reward signals for biological reasoning.
NavitraTechnologies01/navikinase-1.0
by NavitraTechnologies01A from-scratch, decoder-only protein language model for the phosphotransferase superfamily (EC 2.7.-: protein kinases plus sugar/lipid/nucleotide kinases), trained entirely locally on Apple Silicon via MLX β no cloud compute, no fine-tuning of an existing model.
prathmeshadsod/BondShift-Llama-3.3-70B-Instruct
by prathmeshadsodBondShift: Organic Mechanism Reasoning
HuggingFaceBio/Carbon-3B
by HuggingFaceBioTechnical Report π§¬
This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizerβ¦
MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variantβ¦
insilicomedicine/Qwen3-0.6B-Longevity
by insilicomedicine!Format !Task !Params !Type !License
llmithull/HealthGPT-LoRA
by llmithullHealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.
ZeroOneAI/ZEO-Med-2
by ZeroOneAITrinity-Mini-AI-Scientist
This is a QLoRA adapter for query-focused structured extraction from one PubMed title and abstract. It was trained as part of BioEvidence Copilot and targets the repository's versioned ModelEvidenceExtraction JSON Schema.
introvoyz041/DrugGen-2
by introvoyz041# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a diseaseβaware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
Strict automatic scores on the unchanged 1,309-example primary holdout; compare values within each task panel.
RetroAgent is a 4B-parameter LLM agent for multi-step retrosynthesis planning. It decomposes a target molecule into commercially available building blocks by searching over an AND-OR graph of molecules and reactions, driven entirely by tool calls.
trillionlabs/TxGravity-30B-A5B
by trillionlabsTxGravity-30B-A5B is a therapeutics-focused language model fine-tuned from the Gravity-30B-A5B-base. It is trained to predict a broad range of therapeutic properties β small-molecule ADMET, toxicity, drugβtarget interaction, proteinβprotein and peptideβMHC interaction, and more β following theβ¦
alimotahharynia/DrugGen-2
by alimotahharynia# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a diseaseβaware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
CladeTeam/CENO-P-1B
by CladeTeamCENO-P-1B is the multi-species alignment (MSA) post-trained variant of the 1B CENO DNA foundation model, for variant effect prediction (VEP). It carries intraencodingpattern in its config and ships the MSA scoring path (modelingcenop.py), which consumes a per-token seq_idx to score packed MSAβ¦
CladeTeam/CENO-1B-131k
by CladeTeamCENO-1B-131k is the long-context (131k) checkpoint of the 1B CENO DNA foundation model β a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
CladeTeam/CENO-80M-1m
by CladeTeamCENO-80M-1m is the long-context (1M) checkpoint of the 80M CENO DNA foundation model β a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B
by reaperdoesntknowA 1.7B-parameter causal language model distilled from Qwen3-30B-A3B on 6,122 STEM chain-of-thought samples using discrepancy-informed knowledge distillation. The training objective emphasizes proof structure, detects reasoning pivot tokens through token-level divergence dynamics, smoothsβ¦
fableforge-ai/NEXUS-Medical
by fableforge-ai> NEXUS domain specialist for medical Q&A and clinical reasoning β lightweight & uncensored.
!Screenshot 2026-07-05 at 2.33.47 AM
Pippinlitli/evolva-qwen-0.5b-heretic
by PippinlitliHeretic-abliterated version of Qwen/Qwen2.5-0.5B-Instruct for the Evolva drug discovery pipeline.
zsyjsld/Xinghe1-9B
by zsyjsldXinghe1-9B (ζζ Έ) is a specialized large language model fine-tuned for the formalization, computational derivation, and clinical reasoning of Huangdi Neijing. It is based on the Qwen3.5-9B-Instruct architecture and trained using the V3 Double-Purity SFT dataset.
SeongryongJung/Qwen3-4B-Chemistry-SDPO
by SeongryongJungThis repository contains Chemistry fine-tuned Qwen3-4B checkpoints from the local SciKnowEval-style generalization setup.
π Join our LiGHT community. π Check out the MeditronFO blog and MeditronFO preprint. π If you are a clinician join the MOOVE initiative here.
FrenchCastle/sexology-v4
by FrenchCastleSexo-FR is a French-language conversational language model that provides reliable, caring, and evidence-based sexual health information (information en santΓ© sexuelle). It is part of a French public-health initiative whose goal is to make trustworthy sexual-health information more accessible to theβ¦
π Join our LiGHT community. π Check out the MeditronFO blog and MeditronFO preprint. π If you are a clinician join the MOOVE initiative here.
EPFLiGHT/Meditron3-8B
by EPFLiGHTfairydance/molexar-10m-base
by fairydanceMolexar-10M Base is the unconditional base model for Molexar, a unified multimodal molecular foundation model for drug design. It is trained as an autoregressive molecular language model over Fragment-SELFIES, a BRICS-fragment molecular language with validity-preserving decoding andβ¦
fairydance/molexar-10m-omni
by fairydanceMolexar-10M Omni is the universal multi-condition model for Molexar, a unified multimodal molecular foundation model for drug design. It starts from fairydance/molexar-10m-base and is supervised fine-tuned to generate Fragment-SELFIES molecules under scalar molecular-property,β¦
iti-visual-analytics/GRamma-12B
by iti-visual-analyticsGRamma-12B is a 12-billion-parameter instruction-tuned language model specialized for the Greek medical domain. It is built on top of Gemma 3 12B Instruct and adapted through parameter-efficient fine-tuning on a collection of Greek and bilingual medical question-answering data.
QLoRA adapter for Llama-3.1-8B-Instruct, fine-tuned on PubMedQA for yes / no / maybe biomedical question answering (run5).
BioMatrix is a multimodal biological foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture.
nikitaredy/medictron-7B
by nikitaredyA domain-adapted clinical LLM fine-tuned on synthetic Indian medical Q&A records using QLoRA (4-bit quantization) with Unsloth 2x speedup. Built to power the conversational AI layer.
ProtGPT3-MSA is a multiple-sequence, homolog-conditioned autoregressive protein language model. It is part of the ProtGPT3 family, an open-source suite of promptable and aligned protein language models for protein sequence generation.
weblab-LLM-M/AscleLM-1-10B
by weblab-LLM-MBabajaan/KAU-BioMedLLM
by BabajaanKAU-BioMedLLM is a research prototype for source-grounded biomedical variant interpretation. The current public release contains the LoRA adapter and documentation for a guarded report-generation system built around a curated biomedical evidence panel, citation enforcement, and abstention whenβ¦
## Important Notice If you are using GENERator for sequence generation, please ensure that the length of each input sequence is a multiple of 6. This can be achieved by either: 1. Padding the sequence on the left with 'A' (left padding); 2. Truncating the sequence from the left (left truncation).
tokyotech-llm/Medical-GPT-OSS-Swallow-120B
by tokyotech-llmMedical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
tokyotech-llm/Medical-Qwen3-Swallow-30B-A3B
by tokyotech-llmMedical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
gemma4-12b-bioinfo is a fine-tuned Gemma 4 12B instruction model for bioinformatics, genomics, and computational biology question answering.
modelid = "DuanYi/R3LMHepG2"
FINAL-Bench/Darwin-218B-Delphi
by FINAL-Bench> VIDRAFT FINAL-Bench β chemistry-specialized 218B MoE, served via the DELPHI 5-Phase inference cascade.
HealthJudge is a domain-adapted helpfulness evaluator for health-related Community Notes. It is designed to judge whether a note provides helpful context for a potentially misleading social-media post, following the Community Notes helpfulness criteria.