Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

477 of 7,078 resources

Showing 1–50

LoRA adapters fine-tuning Qwen2.5-7B-Instruct to emit a full CIF crystal structure from a prompt of reduced composition + target space-group number. Part of a controlled composition-sweep study (MP-20 : MPTS-52 training ratio at fixed volume/steps).

Active015 hours ago
Python

A DNA annotation model from the Carbon family.

Active1261 day ago
Python

Paper: RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts

Active542 days ago
Python

MedDecider-27B is the most robust mid-size member of the MedDecider family. Given a clinical state (a note, a trial record or any JSON) and a question with 2 to 10 answer options, it returns a calibrated probability for every option in a single scoring pass, without generating text.

Active392 days ago
Python

Vigyan-7B-BioMed-Chem is a domain-specialized LoRA adapter trained on OLMo-2-1124-7B dedicated to organic chemical synthesis, pharmacology, and molecular biology.

Active703 days ago
Python

Typed-decision adapter for Qwen3.5-4B, tuned on clinical question answering, financial news sentiment and structured record-level workflow decisions.

Active413 days ago
Python

Minimal HuggingFace port of Helix-mRNA -- a hybrid Mamba2 / attention language model for full-length mRNA, trained with next-token prediction on single-nucleotide tokens with a codon-start marker.

Active1714 days ago
Python

Longevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-1.7B, was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-1.7B on aging-related multi-omics and clinical data.

Active3585 days ago
Python

A domain-adapted Qwen3.5-9B for aging and longevity biology. L-LLM is the result of continued pretraining + supervised fine-tuning + a reasoning-augmented continuation pass on a multi-domain corpus spanning clinical aging, epigenomics, transcriptomics, proteomics, and genetics.

Active5645 days ago
Python

BioGravity-Inst is a biomedical instruction model from AIGEN Sciences, Inc., developed from the Gravity 30B-A5B family. It is intended for biomedical research in the Biomni A1 environment, including question answering, evidence gathering, computation, and tool-assisted analysis.

Active4045 days ago
Python

A System-1 decision model for biomedicine, pharma and clinical trials. It reads a source, a question and a set of possible answers, and returns a calibrated probability for each answer in one forward pass, with no generated text.

Active2305 days ago
Python

A taxonomy-informed sparse DNA foundation model for microbial genomics.

Active5686 days ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active611 week ago
Python

For a convenient overview and download list, visit our model page for this model.

Active1.9K1 week ago
Python

For a convenient overview and download list, visit our model page for this model.

Active5541 week ago
Python

Byte-level BPE tokenizer for Chargaff, our DNA prediction model. No training, no merges: 1 token per UTF-8 byte.

Active01 week ago
Python

For a convenient overview and download list, visit our model page for this model.

Active3.3K1 week ago
Python

💻 GitHub | 📘 E-SMILES 2.0 Spec | 📄 Report | 🚀 Demo

Active2601 week ago
Python
Active51 week ago
Python
Active82 weeks ago
Python

iona-denoise-50m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 50m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active382 weeks ago
Python

iona-denoise-400m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 400m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active262 weeks ago
Python

iona-denoise-200m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 200m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active372 weeks ago
Python

iona-denoise-100m scores every peak of a tandem mass spectrum (MS/MS) as signal or noise. It is the Iona 100m encoder with a per-peak classification head, fine-tuned for noise-peak detection.

Active382 weeks ago
Python

Three genomic foundation models, packaged together for local inference on Apple silicon.

Active02 weeks ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active272 weeks ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active262 weeks ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active262 weeks ago
Python

Iona is a transformer encoder foundation model for tandem mass spectra (MS/MS). It treats each centroided peak as a token and learns how peaks relate to each other through a per-head attention bias over the signed m/z difference (Δm/z) between every pair of peaks.

Active302 weeks ago
Python

🩺 HuatuoGPT-3-27B 🏠 GitHub | 📄 Paper

Active1202 weeks ago
Python

🩺 HuatuoGPT-3-9B 🏠 GitHub | 📄 Paper

Active2862 weeks ago
Python

HuatuoGPT-3-Grader-8B GitHub | Paper

Active6552 weeks ago
Python

Developed by

Active1882 weeks ago
Python

English | 简体中文

Active6732 weeks ago
Python

Ultra-fast extraction of predefined clinical variables from free-text clinical notes.

Active1042 weeks ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active1.1K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active455.7K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active8.2K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active22K3 weeks ago
Python

ESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…

Active124K3 weeks ago
Python

ESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…

Active181.7K3 weeks ago
Python
Active43 weeks ago
Python

Evo2-7B (Transformers port)

Active1.3K3 weeks ago
Python

Evo2-1B-Base (Transformers port)

Active1.7K3 weeks ago
Python

# Geneformer Geneformer is a foundational transformer model pretrained on a large-scale corpus of human single cell transcriptomes to enable context-aware predictions in settings with limited data in network biology.

Active4.6K4 weeks ago
Python