Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

470 of 7,050 resources

Showing 1–50

Minimal HuggingFace port of Helix-mRNA -- a hybrid Mamba2 / attention language model for full-length mRNA, trained with next-token prediction on single-nucleotide tokens with a codon-start marker.

Active01 day ago
Python

A domain-adapted Qwen3.5-9B for aging and longevity biology. L-LLM is the result of continued pretraining + supervised fine-tuning + a reasoning-augmented continuation pass on a multi-domain corpus spanning clinical aging, epigenomics, transcriptomics, proteomics, and genetics.

Active3301 day ago
Python

BioGravity-Inst is a biomedical instruction model from AIGEN Sciences, Inc., developed from the Gravity 30B-A5B family. It is intended for biomedical research in the Biomni A1 environment, including question answering, evidence gathering, computation, and tool-assisted analysis.

Active3752 days ago
Python

A System-1 decision model for biomedicine, pharma and clinical trials. It reads a source, a question and a set of possible answers, and returns a calibrated probability for each answer in one forward pass, with no generated text.

Active412 days ago
Python

A taxonomy-informed sparse DNA foundation model for microbial genomics.

Active5503 days ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active283 days ago
Python

For a convenient overview and download list, visit our model page for this model.

Active1.9K6 days ago
Python

For a convenient overview and download list, visit our model page for this model.

Active4976 days ago
Python

Byte-level BPE tokenizer for Chargaff, our DNA prediction model. No training, no merges: 1 token per UTF-8 byte.

Active01 week ago
Python

💻 GitHub | 📘 E-SMILES 2.0 Spec | 📄 Report | 🚀 Demo

Active1741 week ago
Python
Active51 week ago
Python
Active81 week ago
Python

iona-denoise-50m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 50m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active351 week ago
Python

iona-denoise-400m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 400m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active261 week ago
Python

iona-denoise-200m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 200m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active341 week ago
Python

iona-denoise-100m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 100m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active351 week ago
Python

Three genomic foundation models, packaged together for local inference on Apple silicon.

Active01 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active241 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active251 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active241 week ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active281 week ago
Python

🩺 HuatuoGPT-3-27B 🏠 GitHub | 📄 Paper

Active1201 week ago
Python

🩺 HuatuoGPT-3-9B 🏠 GitHub | 📄 Paper

Active2861 week ago
Python

HuatuoGPT-3-Grader-8B GitHub | Paper

Active6552 weeks ago
Python

Developed by

Active1882 weeks ago
Python

English | 简体中文

Active6732 weeks ago
Python

Ultra-fast extraction of predefined clinical variables from free-text clinical notes.

Active952 weeks ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active1.1K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active802.1K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active8.8K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active22K3 weeks ago
Python

ESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…

Active208.7K3 weeks ago
Python

ESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…

Active181.7K3 weeks ago
Python
Active43 weeks ago
Python

A domain-adapted small language model specialized for molecular conformation analysis, computational biochemistry, and pharmacology reasoning.

Active963 weeks ago
Python

Evo2-7B (Transformers port)

Active1.3K3 weeks ago
Python

Evo2-1B-Base (Transformers port)

Active1.7K3 weeks ago
Python

# Geneformer Geneformer is a foundational transformer model pretrained on a large-scale corpus of human single cell transcriptomes to enable context-aware predictions in settings with limited data in network biology.

Active4.6K3 weeks ago
Python

> ⚠️ 重要:本仓库的 adapter 历史上因 PeftModel.frompretrained 双重包装导致 key 嵌套错误。 > 旧版本里 PeftModel.frompretrained 加载会"Found missing adapter keys"并静默丢弃全部权重, > 模型实际退化为 base Qwen2.5-3B-Instruct。 > 现在本仓库的 adapter_model.safetensors 已重新打包为标准深度 8(504/504 keys 命中),可被正确加载。 > 验证方式:见 shikunpunk/ask-dao-v0.3 仓库里 "Holdout…

Active783 weeks ago
Python

知识发现机器 —— 从生物医学论文推断「作者没有明说」的开放科学问题

Active03 weeks ago
Python

See the upstream model card for full details, training data and citation.

Active2153 weeks ago
Python

L1-30B-A5B is the Korean-locale medical foundation model from Lunit and Lunit Consortium. It is the 30B member of the L1 family, post-trained directly from Gravity-30B-A5B-Base, a sparse Mixture-of-Experts model developed by Trillion Labs and the Lunit Consortium.

Active3873 weeks ago
Python

💻 Github | 📄 Report | 🚀 Demo

Active2664 weeks ago
Python

For a convenient overview and download list, visit our model page for this model.

Active8031 month ago
Python