Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type(1)
470 of 7,050 resources
Showing 101–150
genzeonplatform/healthcare-brain-ner
by genzeonplatformHealthcare Brain NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated detection and de-identification of Protected Health Information (PHI) and Personally Identifiable Information (PII) in unstructured clinical text.
genzeonplatform/cliniguard-laboratory-ner
by genzeonplatformCliniGuard Laboratory NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of laboratory test results, values, units, reference ranges, and abnormality flags from unstructured clinical text.
genzeonplatform/cliniguard-diagnosis-icd-ner
by genzeonplatformCliniGuard Diagnosis ICD NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of diagnoses, conditions, and support for ICD-10/SNOMED code mapping from unstructured clinical text.
genzeonplatform/cliniguard-medication-ner
by genzeonplatformCliniGuard Medication NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of medication names, dosages, routes, frequencies, and administration details from unstructured clinical text.
CliniGuard Clinical Findings NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of clinical findings, diseases, conditions, anatomical locations, and clinical modifiers from unstructured clinical text.
introvoyz041/DrugGen-2
by introvoyz041# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a disease‑aware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
Strict automatic scores on the unchanged 1,309-example primary holdout; compare values within each task panel.
RetroAgent is a 4B-parameter LLM agent for multi-step retrosynthesis planning. It decomposes a target molecule into commercially available building blocks by searching over an AND-OR graph of molecules and reactions, driven entirely by tool calls.
trillionlabs/TxGravity-30B-A5B
by trillionlabsTxGravity-30B-A5B is a therapeutics-focused language model fine-tuned from the Gravity-30B-A5B-base. It is trained to predict a broad range of therapeutic properties — small-molecule ADMET, toxicity, drug–target interaction, protein–protein and peptide–MHC interaction, and more — following the…
Phsntom/ESMFold2-Fast
by PhsntomESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…
alimotahharynia/DrugGen-2
by alimotahharynia# DrugGen 2: A disease-aware language model for enhancing drug discovery DrugGen-2 is a disease‑aware language model specialized for generating drug-like SMILES structures based on both disease pathways and protein sequence.
CladeTeam/CENO-P-1B
by CladeTeamCENO-P-1B is the multi-species alignment (MSA) post-trained variant of the 1B CENO DNA foundation model, for variant effect prediction (VEP). It carries intraencodingpattern in its config and ships the MSA scoring path (modelingcenop.py), which consumes a per-token seq_idx to score packed MSA…
CladeTeam/CENO-1B-131k
by CladeTeamCENO-1B-131k is the long-context (131k) checkpoint of the 1B CENO DNA foundation model — a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
CladeTeam/CENO-80M-1m
by CladeTeamCENO-80M-1m is the long-context (1M) checkpoint of the 80M CENO DNA foundation model — a causal language model over genomic sequence built on a Nemotron-H Mamba / Attention / Mixture-of-Experts hybrid backbone (no MSA inputs).
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B
by reaperdoesntknowA 1.7B-parameter causal language model distilled from Qwen3-30B-A3B on 6,122 STEM chain-of-thought samples using discrepancy-informed knowledge distillation. The training objective emphasizes proof structure, detects reasoning pivot tokens through token-level divergence dynamics, smooths…
Fine-tuned ESM-2 650M with LoRA for predicting protein subcellular localization (10 classes).
phenobase/phenovisionL
by phenobasePhenoVisionL is a Vision Transformer (ViT-Large) model fine-tuned to detect leaf phenological states in plant photographs: green leaves, colored (senescent) leaves, and breaking leaf buds. It was trained on 165,988 iNaturalist records of deciduous woody plants using a two-stage semi-supervised…
phenobase/phenovision
by phenobasePhenoVision is a Vision Transformer (ViT-Large) model fine-tuned to detect flowers and fruits in plant photographs. It was trained on 1.5 million human-annotated iNaturalist images and has been used to generate over 30 million new phenology records across 119,000+ plant species, vastly expanding…
fableforge-ai/NEXUS-Medical
by fableforge-ai> NEXUS domain specialist for medical Q&A and clinical reasoning — lightweight & uncensored.
Clinical-Reasoning-Hub/pentabrid-27b
by Clinical-Reasoning-Hublowdown-labs/fela-genomics
by lowdown-labslowdown-labs/fela-chemistry
by lowdown-labs!Screenshot 2026-07-05 at 2.33.47 AM
This repository contains LoRA finetunes of DiffusionGemma (image-conditioned discrete-diffusion LLM) for radiology visual question answering, each paired with an autoregressive Gemma-4 finetune as a controlled baseline. It corresponds to the paper Discrete Diffusion Language Models for Interactive…
doctolib-lab/doctobert-fr-base
by doctolib-lab🤗 Blog | 📄 Paper | 💻 Code | 🌐 FineMed | 🩺 DoctoBERT
Pippinlitli/evolva-qwen-0.5b-heretic
by PippinlitliHeretic-abliterated version of Qwen/Qwen2.5-0.5B-Instruct for the Evolva drug discovery pipeline.
HantaBERT/HantaBERT
by HantaBERTHantaBERT fine-tunes DNABERT-2 on hantavirus RNA sequences for three simultaneous classification tasks: species/lineage, host, and geographic origin. A single forward pass produces predictions for all three tasks along with a 768-dimensional embedding suitable for phylogenetic visualization.
mradermacher/CellHermes-v1.0-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
zsyjsld/Xinghe1-9B
by zsyjsldXinghe1-9B (杏核) is a specialized large language model fine-tuned for the formalization, computational derivation, and clinical reasoning of Huangdi Neijing. It is based on the Qwen3.5-9B-Instruct architecture and trained using the V3 Double-Purity SFT dataset.
EthanGao123/CellHermes-v1.0
by EthanGao123# Overview This is the CellHermes model, based on the LLaMA-3.1-8B-instruct architecture developed by Meta, fine-tuned using single-cell RNA sequencing (scRNA-seq) datasets from CellxGene and PPI network from BioGRID. CellHermes is an innovative framework for adapting existing large language models…
SeongryongJung/Qwen3-4B-Chemistry-SDPO
by SeongryongJungThis repository contains Chemistry fine-tuned Qwen3-4B checkpoints from the local SciKnowEval-style generalization setup.
PlantGeneAnn is a plant genome foundation model that enables the prediction of various plant genomic elements at single-nucleotide resolution. The model is built upon the PlantBiMoE architecture with a 1D U-Net segmentation head, specifically designed for automated plant genome annotation.
👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here.
FrenchCastle/sexology-v4
by FrenchCastleSexo-FR is a French-language conversational language model that provides reliable, caring, and evidence-based sexual health information (information en santé sexuelle). It is part of a French public-health initiative whose goal is to make trustworthy sexual-health information more accessible to the…
👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here.
trillionlabs/Gravity-bio-16B-A3B
by trillionlabsGravity-bio-16B-A3B is a biology-focus midtrained model derived from Gravity-16B-A3B-Base. It uses the same sparse Mixture-of-Experts (MoE) architecture and tokenizer as Gravity-16B-A3B-Base, with additional midtraining for biological understanding on TheBioCollection corpus.
EPFLiGHT/Meditron3-8B
by EPFLiGHTraidium/Jolia
by raidiumJolia is a 3D CT foundation model that encodes images into vector representations program. It encodes a whole 3D CT volume into:
fairydance/molexar-10m-base
by fairydanceMolexar-10M Base is the unconditional base model for Molexar, a unified multimodal molecular foundation model for drug design. It is trained as an autoregressive molecular language model over Fragment-SELFIES, a BRICS-fragment molecular language with validity-preserving decoding and…
fairydance/molexar-10m-omni
by fairydanceMolexar-10M Omni is the universal multi-condition model for Molexar, a unified multimodal molecular foundation model for drug design. It starts from fairydance/molexar-10m-base and is supervised fine-tuned to generate Fragment-SELFIES molecules under scalar molecular-property,…
iti-visual-analytics/GRamma-12B
by iti-visual-analyticsGRamma-12B is a 12-billion-parameter instruction-tuned language model specialized for the Greek medical domain. It is built on top of Gemma 3 12B Instruct and adapted through parameter-efficient fine-tuning on a collection of Greek and bilingual medical question-answering data.
Full weight-level fine-tuning of InstaDeepAI/nucleotide-transformer-v2-50m-multi-species for binary DNA sequence classification on two GenomicBenchmarks tasks. All parameters are updated rather than using LoRA or a frozen backbone, with a leakage-free train/validation/test protocol and multi-seed…
QLoRA adapter for Llama-3.1-8B-Instruct, fine-tuned on PubMedQA for yes / no / maybe biomedical question answering (run5).
BioMatrix is a multimodal biological foundation model that natively integrates 1D sequences, 3D structures, and natural language for both molecules and proteins within a single decoder-only architecture.
nikitaredy/medictron-7B
by nikitaredyA domain-adapted clinical LLM fine-tuned on synthetic Indian medical Q&A records using QLoRA (4-bit quantization) with Unsloth 2x speedup. Built to power the conversational AI layer.
ProtGPT3-MSA is a multiple-sequence, homolog-conditioned autoregressive protein language model. It is part of the ProtGPT3 family, an open-source suite of promptable and aligned protein language models for protein sequence generation.