Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source(1)
Type
887 of 7,078 resources
Showing 1β50
shehrozashoaib/LLM_Crystal_CIF
by shehrozashoaibLoRA adapters fine-tuning Qwen2.5-7B-Instruct to emit a full CIF crystal structure from a prompt of reduced composition + target space-group number. Part of a controlled composition-sweep study (MP-20 : MPTS-52 training ratio at fixed volume/steps).
HuggingFaceBio/Carbon-A-1.2B
by HuggingFaceBioA DNA annotation model from the Carbon family.
pantoniadis/RAGenome
by pantoniadisPaper: RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts
MedDecider-27B is the most robust mid-size member of the MedDecider family. Given a clinical state (a note, a trial record or any JSON) and a question with 2 to 10 answer options, it returns a calibrated probability for every option in a single scoring pass, without generating text.
StanfordShahLab/motor-t-base
by StanfordShahLabStanfordShahLab/clmbr-t-base
by StanfordShahLabshreyansh12183/olmo2-7b-biomed-chem
by shreyansh12183Vigyan-7B-BioMed-Chem is a domain-specialized LoRA adapter trained on OLMo-2-1124-7B dedicated to organic chemical synthesis, pharmacology, and molecular biology.
helloimsaif/medjev-4b-lora
by helloimsaifTyped-decision adapter for Qwen3.5-4B, tuned on clinical question answering, financial news sentiment and structured record-level workflow decisions.
atom-jepa/atom-jepa
by atom-jepaPretrained encoders from Atom-JEPA: Joint-Embedding Predictive Architecture for 3D Atomistic Systems.
Taykhoom/Helix-mRNA
by TaykhoomMinimal HuggingFace port of Helix-mRNA -- a hybrid Mamba2 / attention language model for full-length mRNA, trained with next-token prediction on single-nucleotide tokens with a codon-start marker.
insilicomedicine/Qwen3-1.7B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-1.7B, was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-1.7B on aging-related multi-omics and clinical data.
insilicomedicine/longevity-llm
by insilicomedicineA domain-adapted Qwen3.5-9B for aging and longevity biology. L-LLM is the result of continued pretraining + supervised fine-tuning + a reasoning-augmented continuation pass on a multi-domain corpus spanning clinical aging, epigenomics, transcriptomics, proteomics, and genetics.
aigensciences/BioGravity-Inst
by aigensciencesBioGravity-Inst is a biomedical instruction model from AIGEN Sciences, Inc., developed from the Gravity 30B-A5B family. It is intended for biomedical research in the Biomni A1 environment, including question answering, evidence gathering, computation, and tool-assisted analysis.
lighteternal/biodecision-v2-4b
by lighteternalA System-1 decision model for biomedicine, pharma and clinical trials. It reads a source, a question and a set of possible answers, and returns a calibrated probability for each answer in one forward pass, with no generated text.
athanzli/MicroGlot
by athanzliA taxonomy-informed sparse DNA foundation model for microbial genomics.
drzo/ESMC-6B
by drzoESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into proteinβ¦
Madrimed1.2 2B GGUF Optimized GGUF Quantization Tiers for MadriMed 1.2 (2B Medical Mulitmodel)
asgersvenning/MAMBO-v3
by asgersvenningNemo (MAMBO_v3) identifies adult moths and butterflies in photographs, predicting 12,632 species, 4,476 genera and 104 families. Predictions use GBIF taxon IDs.
For a convenient overview and download list, visit our model page for this model.
For a convenient overview and download list, visit our model page for this model.
madrisight/MadriMed-VL-2B
by madrisightA 2B-parameter multimodal medical vision-language model, trained for medical image understanding, radiology report assistance, clinical visual question answering, and medical text reasoning.
Aquiles-ai/Chargaff-Tokenizer
by Aquiles-aiByte-level BPE tokenizer for Chargaff, our DNA prediction model. No training, no merges: 1 token per UTF-8 byte.
marvinsxtr/MapPFN
by marvinsxtrPre-trained and fine-tuned checkpoints for MapPFN: Learning Causal Perturbation Maps in Context (Sextro et al., 2026).
For a convenient overview and download list, visit our model page for this model.
joelinator/dflow-novo-model
by joelinatorDFlowNovo is a state-of-the-art de novo peptide sequencing system powered by Continuous-Time Markov Chain Discrete Flow Matching (CTMC-DFM). By treating peptide sequencing as continuous probability flows over discrete amino acid states and integrating dynamic programming (KnapsackDP) reachabilityβ¦
autoencodix/Ontix-Dim16-KimiK3
by autoencodixAn Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with Kimi K3, making the embedding directly interpretable in terms of biological processes.
autoencodix/Ontix-Dim16-GPT5-6
by autoencodixAn Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with GPT-5.6 TerraPro, making the embedding directly interpretable in terms of biological processes.
An Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with Claude Opus 5, making the embedding directly interpretable in terms of biological processes.
UniParser/MolParser-Mobile-V2
by UniParserπ» GitHub | π E-SMILES 2.0 Spec | π Report | π Demo
Haocheng1/CrystAF
by Haocheng1Weights for Where Should Physics Enter a Molecular Crystal Generator? (Haocheng Tang, Junmei Wang, Wengong Jin).
Soilytix/LOAM-624M
by SoilytixContrastive LEarning with Soft Targets from TCRdist. Checkpoint SCEPTR6LACsoft800k20ep_bs1024.
Gaolaboratory/iona-denoise-50m
by Gaolaboratoryiona-denoise-50m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 50m encoder with a per-peak classification head, fine-tuned for noise-peak detection.
Gaolaboratory/iona-denoise-400m
by Gaolaboratoryiona-denoise-400m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 400m encoder with a per-peak classification head, fine-tuned for noise-peak detection.
Gaolaboratory/iona-denoise-200m
by Gaolaboratoryiona-denoise-200m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 200m encoder with a per-peak classification head, fine-tuned for noise-peak detection.
Gaolaboratory/iona-denoise-100m
by Gaolaboratoryiona-denoise-100m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 100m encoder with a per-peak classification head, fine-tuned for noise-peak detection.
Three genomic foundation models, packaged together for local inference on Apple silicon.
Gaolaboratory/iona-base-400m
by GaolaboratoryIona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Ξm/z) between every pair of peaks.
Gaolaboratory/iona-base-200m
by GaolaboratoryIona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Ξm/z) between every pair of peaks.
Gaolaboratory/iona-base-100m
by GaolaboratoryIona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Ξm/z) between every pair of peaks.
Gaolaboratory/iona-base-50m
by GaolaboratoryIona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Ξm/z) between every pair of peaks.
gaozijun/CELLO
by gaozijunFreedomIntelligence/HuatuoGPT-3-27B
by FreedomIntelligenceπ©Ί HuatuoGPT-3-27B π GitHub | π Paper
FreedomIntelligence/HuatuoGPT-3-9B
by FreedomIntelligenceπ©Ί HuatuoGPT-3-9B π GitHub | π Paper
faerte/neural_paw_dft
by faerteTrained weights for the paper Complete Neural Electronic Initialization Accelerates Materials DFT (arXiv:2609.21759).
FreedomIntelligence/HuatuoGPT-3-Grader-8B
by FreedomIntelligenceHuatuoGPT-3-Grader-8B GitHub | Paper
Kentucky-Open-Science/KOS-V5-Instruct
by Kentucky-Open-ScienceDeveloped by
migi-null/DeepGPS-3D
by migi-nullPretrained weights for the baseline DeepGPS-3D model: a conditional 3D denoising diffusion model that predicts a protein's 3D subcellular localization volume from a matched nuclear 3D volume and the protein's ESM2 sequence embedding.