Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

19 of 7,064 resources

Phylogeny-aware genomic language model trained on whole-genome alignments across multiple evolutionary timescales, predicting functional constraints and variant effects for human, mouse, chicken, fly, worm, and Arabidopsis genomes (344+ stars, MIT License)

Active3674 weeks ago
Jupyter Notebook
MIT

Large transformer-based single-cell foundation model pretrained on 50 million cells for robust gene network inference, expression denoising, cell embedding, and zero-shot label prediction, leveraging ESM2 protein embeddings and bidirectional transformer architecture (Cantini Lab, 148+ stars, GPL-3.0)

Active1621 month ago
Jupyter Notebook
GPL-3.0

Arc Institute's 40B-parameter genome foundation model trained on 9 trillion nucleotides from all domains of life, supporting 1M base pair context for generalist DNA/RNA/protein prediction and design (Nature 2026)

Active4.2K3 months ago
Jupyter Notebook
Apache-2.0

Gene expression prediction

Active15.2K3 months ago
Jupyter Notebook
Apache-2.0

Transformer foundation model for tandem mass spectrometry (MS/MS) self-supervised on millions of unannotated spectra from the GeMS dataset via masked peak prediction and chromatographic retention-order objectives, producing 1024-dimensional molecular representations; achieves SOTA on spectral similarity, chemical property, and molecular fingerprint prediction, and powers the DreaMS Atlas annotating 201M+ MS/MS spectra for metabolomics and natural product discovery (Pluskal Lab, IOCB Prague & MIT, 211+ stars, MIT License)

Active2124 months ago
Jupyter Notebook
MIT

First architecture deeply integrating a DNA foundation model with an LLM for multimodal biological reasoning, achieving 98% accuracy on KEGG disease pathway prediction and 15%+ average gains on variant effect prediction with interpretable step-by-step reasoning traces (bowang-lab, 390+ stars)

Active4024 months ago
Jupyter Notebook
Apache-2.0

Multimodal AI bridging transcriptomics data and natural language, enabling intuitive chat-based exploration and analysis of single-cell RNA-seq datasets through conversational interaction without coding; fine-tuned Mistral 7B LLaVA model emulating biologist-bioinformatician discussions (207+ stars, GPL-3.0)

Active2194 months ago
Jupyter Notebook
GPL-3.0

Single-cell analysis with transformers

Active1.6K5 months ago
Jupyter Notebook
MIT

Arc Institute's single-cell foundation model enabling in-context learning at inference time via a novel tabular attention architecture, trained on 150M uniformly-preprocessed cells for generalizing biological effects and generating unseen cell profiles in novel contexts (2025)

Active1625 months ago
Jupyter Notebook
NOASSERTION

Foundation models for genomics and transcriptomics pretrained on 3,000+ human genomes and 850+ diverse species, enabling chromatin accessibility prediction, splice site detection, and promoter classification across multiple model scales (InstaDeep, NVIDIA & TUM, Nature Methods 2023)

Idle9197 months ago
Jupyter Notebook
NOASSERTION

Generative AI framework for inverse design of 3D RNA structure and function using geometric deep learning, learning design rules from 3D structures to capture complex tertiary interactions (pseudoknots, non-canonical base pairs) with expert-level accuracy for designing functional RNAs including aptamers and ribozymes (bioRxiv 2025)

Idle3209 months ago
Jupyter Notebook
MIT

Foundation model jointly trained on single-cell and spatial transcriptomics data, enabling unified representation learning across cellular and tissue spatial contexts for cell type prediction, spatial domain inference, and cross-modal integration (theislab, bioRxiv 2024, 164+ stars)

Idle17510 months ago
Jupyter Notebook
BSD-3-Clause

100M-parameter foundation model pretrained on 50M+ human single-cell transcriptomes covering ~20,000 genes, achieving SOTA on gene expression enhancement, drug response and perturbation prediction (Nature Methods 2024)

Idle43110 months ago
Jupyter Notebook
Apache-2.0

Family of generative single-cell foundation models (TF-Metazoa, TF-Exemplar, TF-Sapiens) jointly modeling genes and their expression levels via expression-aware autoregressive transformers, trained on up to 112M cells across 12 species spanning 1.53 billion years of evolution; achieves robust zero-shot cell type classification across species, disease state identification in human cells, and prediction of cell-type-specific transcription factors and gene-gene regulatory relationships, pip-installable with pretrained weights (CZI, 166+ stars, MIT License)

Idle16611 months ago
Jupyter Notebook
MIT

Teaching Large Language Models the Language of Biology through single-cell transcriptomics (ICML 2024)

Idle87811 months ago
Jupyter Notebook
Apache-2.0

End-to-end deep learning approach for RNA tertiary structure prediction with a flexible nucleobase center representation, achieving ~7 Å C1' RMSD across test RNAs and predicting ~545,000 structures covering 2,200+ RNA families (Kihara Lab, Purdue University, 50+ stars)

Idle5411 months ago
Jupyter Notebook
GPL-3.0

Pre-trained large generative model translating single-cell transcriptomes to proteomes in an alignment-free manner, generating absent protein abundance data for CITE-seq, spatial CITE-seq, REAP-seq, and NEAT-seq across tissues and diseases; offers three model variants pretrained on 2M human cells, 160K PBMCs, or 18K bulk samples (Tencent AI Lab Healthcare, 96+ stars)

Idle991 year ago
Jupyter Notebook

RNA foundation model trained on millions of RNA sequences for generalist RNA sequence understanding, enabling downstream structure prediction, function annotation, and representation learning for non-coding RNAs (ml4bio, 372+ stars)

Idle3891 year ago
Jupyter Notebook
MIT

Generative pre-training for genomics

Stale3242 years ago
Jupyter Notebook