Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

869 of 7,050 resources

Showing 1–50

Minimal HuggingFace port of Helix-mRNA -- a hybrid Mamba2 / attention language model for full-length mRNA, trained with next-token prediction on single-nucleotide tokens with a codon-start marker.

Active01 day ago
Python

A domain-adapted Qwen3.5-9B for aging and longevity biology. L-LLM is the result of continued pretraining + supervised fine-tuning + a reasoning-augmented continuation pass on a multi-domain corpus spanning clinical aging, epigenomics, transcriptomics, proteomics, and genetics.

Active3301 day ago
Python

BioGravity-Inst is a biomedical instruction model from AIGEN Sciences, Inc., developed from the Gravity 30B-A5B family. It is intended for biomedical research in the Biomni A1 environment, including question answering, evidence gathering, computation, and tool-assisted analysis.

Active3752 days ago
Python

A System-1 decision model for biomedicine, pharma and clinical trials. It reads a source, a question and a set of possible answers, and returns a calibrated probability for each answer in one forward pass, with no generated text.

Active412 days ago
Python

A taxonomy-informed sparse DNA foundation model for microbial genomics.

Active5503 days ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active283 days ago
Python

Nemo (MAMBO_v3) identifies adult moths and butterflies in photographs, predicting 12,632 species, 4,476 genera and 104 families. Predictions use GBIF taxon IDs.

Active04 days ago

For a convenient overview and download list, visit our model page for this model.

Active1.9K6 days ago
Python

For a convenient overview and download list, visit our model page for this model.

Active4976 days ago
Python

Byte-level BPE tokenizer for Chargaff, our DNA prediction model. No training, no merges: 1 token per UTF-8 byte.

Active01 week ago
Python

Pre-trained and fine-tuned checkpoints for MapPFN: Learning Causal Perturbation Maps in Context (Sextro et al., 2026).

Active01 week ago

DFlowNovo is a state-of-the-art de novo peptide sequencing system powered by Continuous-Time Markov Chain Discrete Flow Matching (CTMC-DFM). By treating peptide sequencing as continuous probability flows over discrete amino acid states and integrating dynamic programming (KnapsackDP) reachability…

Active01 week ago

An Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with Kimi K3, making the embedding directly interpretable in terms of biological processes.

Active01 week ago

An Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with GPT-5.6 TerraPro, making the embedding directly interpretable in terms of biological processes.

Active01 week ago

An Ontix autoencoder with an explainable, 16-dimensional latent space, trained on single-cell RNA-seq data. Each latent dimension is constrained by a gene ontology term generated with Claude Opus 5, making the embedding directly interpretable in terms of biological processes.

Active01 week ago

💻 GitHub | 📘 E-SMILES 2.0 Spec | 📄 Report | 🚀 Demo

Active1741 week ago
Python

Weights for Where Should Physics Enter a Molecular Crystal Generator? (Haocheng Tang, Junmei Wang, Wengong Jin).

Active01 week ago
Active51 week ago
Python
Active81 week ago
Python

Contrastive LEarning with Soft Targets from TCRdist. Checkpoint SCEPTR6LACsoft800k20ep_bs1024.

Active121 week ago

iona-denoise-50m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 50m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active351 week ago
Python

iona-denoise-400m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 400m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active261 week ago
Python

iona-denoise-200m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 200m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active341 week ago
Python

iona-denoise-100m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 100m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active351 week ago
Python

Three genomic foundation models, packaged together for local inference on Apple silicon.

Active01 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active241 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active251 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active241 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active281 week ago
Python
Active91 week ago

🩺 HuatuoGPT-3-27B 🏠 GitHub | 📄 Paper

Active1201 week ago
Python

🩺 HuatuoGPT-3-9B 🏠 GitHub | 📄 Paper

Active2861 week ago
Python

Trained weights for the paper Complete Neural Electronic Initialization Accelerates Materials DFT (arXiv:2609.21759).

Active01 week ago

HuatuoGPT-3-Grader-8B GitHub | Paper

Active6551 week ago
Python

Developed by

Active1882 weeks ago
Python

Pretrained weights for the baseline DeepGPS-3D model: a conditional 3D denoising diffusion model that predicts a protein's 3D subcellular localization volume from a matched nuclear 3D volume and the protein's ESM2 sequence embedding.

Active02 weeks ago

Alibaba-DAMO-Academy/RADAR

by Alibaba-DAMO-Academy

# RADAR: An Expert-Level Generalist AI for Abdominal CT Diagnosis

Active02 weeks ago

English | 简体中文

Active6732 weeks ago
Python

Ultra-fast extraction of predefined clinical variables from free-text clinical notes.

Active952 weeks ago
Python

radar-generalist/RADAR

by radar-generalist

# RADAR: An Expert-Level Generalist AI for Abdominal CT Diagnosis

Active02 weeks ago

Fx-Bio-0913 is a biomedical reasoning large language model post-trained on DeepSeek-V4-Flash, developed by The Endless Frontier lab. It is specialized for biological and biomedical research tasks — including gene-function puzzles, experimental reasoning, and multi-step evidence integration —…

Active4512 weeks ago

ChatterjeeLab/PIVOT

by ChatterjeeLab

!PIVOT overview

Active02 weeks ago

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active1.1K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active802.1K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active8.8K3 weeks ago
Python