Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source(1)
Type
377 of 6,573 resources
Showing 51β100
RetroAgent is a 4B-parameter LLM agent for multi-step retrosynthesis planning. It decomposes a target molecule into commercially available building blocks by searching over an AND-OR graph of molecules and reactions, driven entirely by tool calls.
trillionlabs/TxGravity-30B-A5B
by trillionlabsTxGravity-30B-A5B is a therapeutics-focused language model fine-tuned from the Gravity-30B-A5B-base. It is trained to predict a broad range of therapeutic properties β small-molecule ADMET, toxicity, drugβtarget interaction, proteinβprotein and peptideβMHC interaction, and more β following theβ¦
alimotahharynia/DrugGen-2
by alimotahharynia# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a diseaseβaware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
CladeTeam/CENO-P-1B
by CladeTeamCENO-P-1B is the multi-species alignment (MSA) post-trained variant of the 1B CENO DNA foundation model, for variant effect prediction (VEP). It carries intraencodingpattern in its config and ships the MSA scoring path (modelingcenop.py), which consumes a per-token seq_idx to score packed MSAβ¦
CladeTeam/CENO-1B-131k
by CladeTeamCENO-1B-131k is the long-context (131k) checkpoint of the 1B CENO DNA foundation model β a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
CladeTeam/CENO-80M-1m
by CladeTeamCENO-80M-1m is the long-context (1M) checkpoint of the 80M CENO DNA foundation model β a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
UniParser/MolParser-Mobile
by UniParserπ» Github | π Report (Coming soon...) | π Demo
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B
by reaperdoesntknowA 1.7B-parameter causal language model distilled from Qwen3-30B-A3B on 6,122 STEM chain-of-thought samples using discrepancy-informed knowledge distillation. The training objective emphasizes proof structure, detects reasoning pivot tokens through token-level divergence dynamics, smoothsβ¦
Fine-tuned ESM-2 650M with LoRA for predicting protein subcellular localization (10 classes).
phenobase/phenovisionL
by phenobasePhenoVisionL is a Vision Transformer (ViT-Large) model fine-tuned to detect leaf phenological states in plant photographs: green leaves, colored (senescent) leaves, and breaking leaf buds. It was trained on 165,988 iNaturalist records of deciduous woody plants using a two-stage semi-supervisedβ¦
phenobase/phenovision
by phenobasePhenoVision is a Vision Transformer (ViT-Large) model fine-tuned to detect flowers and fruits in plant photographs. It was trained on 1.5 million human-annotated iNaturalist images and has been used to generate over 30 million new phenology records across 119,000+ plant species, vastly expandingβ¦
fableforge-ai/NEXUS-Medical
by fableforge-ai> NEXUS domain specialist for medical Q&A and clinical reasoning β lightweight & uncensored.
Clinical-Reasoning-Hub/pentabrid-27b
by Clinical-Reasoning-Hublowdown-labs/fela-genomics
by lowdown-labslowdown-labs/fela-chemistry
by lowdown-labs!Screenshot 2026-07-05 at 2.33.47 AM
This repository contains LoRA finetunes of DiffusionGemma (image-conditioned discrete-diffusion LLM) for radiology visual question answering, each paired with an autoregressive Gemma-4 finetune as a controlled baseline. It corresponds to the paper Discrete Diffusion Language Models for Interactiveβ¦
doctolib-lab/doctobert-fr-base
by doctolib-labπ€ Blog | π Paper | π» Code | π FineMed | π©Ί DoctoBERT
Pippinlitli/evolva-qwen-0.5b-heretic
by PippinlitliHeretic-abliterated version of Qwen/Qwen2.5-0.5B-Instruct for the Evolva drug discovery pipeline.
HantaBERT/HantaBERT
by HantaBERTHantaBERT fine-tunes DNABERT-2 on hantavirus RNA sequences for three simultaneous classification tasks: species/lineage, host, and geographic origin. A single forward pass produces predictions for all three tasks along with a 768-dimensional embedding suitable for phylogenetic visualization.
mradermacher/CellHermes-v1.0-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
zsyjsld/Xinghe1-9B
by zsyjsldXinghe1-9B (ζζ Έ) is a specialized large language model fine-tuned for the formalization, computational derivation, and clinical reasoning of Huangdi Neijing. It is based on the Qwen3.5-9B-Instruct architecture and trained using the V3 Double-Purity SFT dataset.
EthanGao123/CellHermes-v1.0
by EthanGao123# Overview This is the CellHermes model, based on the LLaMA-3.1-8B-instruct architecture developed by Meta, fine-tuned using single-cell RNA sequencing (scRNA-seq) datasets from CellxGene and PPI network from BioGRID. CellHermes is an innovative framework for adapting existing large language modelsβ¦
SeongryongJung/Qwen3-4B-Chemistry-SDPO
by SeongryongJungThis repository contains Chemistry fine-tuned Qwen3-4B checkpoints from the local SciKnowEval-style generalization setup.
PlantGeneAnn is a plant genome foundation model that enables the prediction of various plant genomic elements at single-nucleotide resolution. The model is built upon the PlantBiMoE architecture with a 1D U-Net segmentation head, specifically designed for automated plant genome annotation.
π Join our LiGHT community. π Check out the MeditronFO blog and MeditronFO preprint. π If you are a clinician join the MOOVE initiative here.
FrenchCastle/sexology-v4
by FrenchCastleSexo-FR is a French-language conversational language model that provides reliable, caring, and evidence-based sexual health information (information en santΓ© sexuelle). It is part of a French public-health initiative whose goal is to make trustworthy sexual-health information more accessible to theβ¦
π Join our LiGHT community. π Check out the MeditronFO blog and MeditronFO preprint. π If you are a clinician join the MOOVE initiative here.
EPFLiGHT/Meditron3-8B
by EPFLiGHTraidium/Jolia
by raidiumJolia is a 3D CT foundation model that encodes images into vector representations program. It encodes a whole 3D CT volume into:
fairydance/molexar-10m-base
by fairydanceMolexar-10M Base is the unconditional base model for Molexar, a unified multimodal molecular foundation model for drug design. It is trained as an autoregressive molecular language model over Fragment-SELFIES, a BRICS-fragment molecular language with validity-preserving decoding andβ¦
fairydance/molexar-10m-omni
by fairydanceMolexar-10M Omni is the universal multi-condition model for Molexar, a unified multimodal molecular foundation model for drug design. It starts from fairydance/molexar-10m-base and is supervised fine-tuned to generate Fragment-SELFIES molecules under scalar molecular-property,β¦
iti-visual-analytics/GRamma-12B
by iti-visual-analyticsGRamma-12B is a 12-billion-parameter instruction-tuned language model specialized for the Greek medical domain. It is built on top of Gemma 3 12B Instruct and adapted through parameter-efficient fine-tuning on a collection of Greek and bilingual medical question-answering data.
Full weight-level fine-tuning of InstaDeepAI/nucleotide-transformer-v2-50m-multi-species for binary DNA sequence classification on two GenomicBenchmarks tasks. All parameters are updated rather than using LoRA or a frozen backbone, with a leakage-free train/validation/test protocol and multi-seedβ¦
QLoRA adapter for Llama-3.1-8B-Instruct, fine-tuned on PubMedQA for yes / no / maybe biomedical question answering (run5).
BioMatrix is a multimodal biological foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture.
nikitaredy/medictron-7B
by nikitaredyA domain-adapted clinical LLM fine-tuned on synthetic Indian medical Q&A records using QLoRA (4-bit quantization) with Unsloth 2x speedup. Built to power the conversational AI layer.
ProtGPT3-MSA is a multiple-sequence, homolog-conditioned autoregressive protein language model. It is part of the ProtGPT3 family, an open-source suite of promptable and aligned protein language models for protein sequence generation.
weblab-LLM-M/AscleLM-1-10B
by weblab-LLM-MBabajaan/KAU-BioMedLLM
by BabajaanKAU-BioMedLLM is a research prototype for source-grounded biomedical variant interpretation. The current public release contains the LoRA adapter and documentation for a guarded report-generation system built around a curated biomedical evidence panel, citation enforcement, and abstention whenβ¦
divyeshkamalanaban/nucleotide-transformer-2.5b-multi-species-NF4-Q4
by divyeshkamalanabanThis model is an NF4 (Normal Float 4-bit) quantized version of the base model InstaDeepAI/nucleotide-transformer-2.5b-multi-species. The checkpoint was quantized using the BitsAndBytes library with double quantization enabled and BF16 computation.
## Important Notice If you are using GENERator for sequence generation, please ensure that the length of each input sequence is a multiple of 6. This can be achieved by either: 1. Padding the sequence on the left with 'A' (left padding); 2. Truncating the sequence from the left (left truncation).
tokyotech-llm/Medical-GPT-OSS-Swallow-120B
by tokyotech-llmMedical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
tokyotech-llm/Medical-Qwen3-Swallow-30B-A3B
by tokyotech-llmMedical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
mradermacher/Phi-4-Instruct-Bioaligned-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
gemma4-12b-bioinfo is a fine-tuned Gemma 4 12B instruction model for bioinformatics, genomics, and computational biology question answering.
modelid = "DuanYi/R3LMHepG2"