Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

69 of 7,078 resources

Showing 51–69

State-specific protein-ligand complex structure prediction with a multi-scale deep generative model, enabling conformational state-aware modeling of molecular interactions (329+ stars, 2024)

Idle3361 year ago
Jupyter Notebook
BSD-3-Clause

AI-assisted mutation nomination approach optimizing protein function by integrating structural and evolutionary constraints into protein inverse folding models, compatible with ProteinMPNN, LigandMPNN, ESM-IF1, and SaProt (Chinese Academy of Sciences, 359+ stars)

Idle1.1K1 year ago
Jupyter Notebook
NOASSERTION

Strongest open-source automated theorem prover in Lean 4, 8B model matches DeepSeek-Prover-V2-671B at 84.6% MiniF2F, 32B model achieves 90.4% with self-correction, using scaffolded data synthesis and verifier-guided proof refinement (Princeton, 2025)

Idle1891 year ago
Jupyter Notebook

Pre-trained large generative model translating single-cell transcriptomes to proteomes in an alignment-free manner, generating absent protein abundance data for CITE-seq, spatial CITE-seq, REAP-seq, and NEAT-seq across tissues and diseases; offers three model variants pretrained on 2M human cells, 160K PBMCs, or 18K bulk samples (Tencent AI Lab Healthcare, 96+ stars)

Idle991 year ago
Jupyter Notebook

Therapeutics Data Commons: 66 AI-ready datasets across 22 drug discovery tasks with 29 leaderboards, covering target identification, molecular generation, ADMET prediction, and clinical trial outcomes (Harvard MIMS, NeurIPS 2021/2024)

Idle1.3K1 year ago
Jupyter Notebook
MIT

RNA foundation model trained on millions of RNA sequences for generalist RNA sequence understanding, enabling downstream structure prediction, function annotation, and representation learning for non-coding RNAs (ml4bio, 372+ stars)

Idle3891 year ago
Jupyter Notebook
MIT

State-of-the-art pretrained language models for proteins trained on thousands of GPUs and Google TPUs using Transformer architectures, enabling protein property prediction, feature extraction, and transfer learning across diverse downstream tasks (1.3K+ stars, MIT, 2020-2026)

Idle1.3K1 year ago
Jupyter Notebook
MIT

First end-to-end data-driven weather prediction system learning directly from raw, heterogeneous Earth observations rather than physics-based reanalysis, producing both global gridded and arbitrary station forecasts; releases model weights, training code, and an ML-ready observational dataset (2007-2019) for building future end-to-end weather models (221+ stars, CC0-1.0)

Idle2211 year ago
Jupyter Notebook
CC0-1.0

Universal medical image segmentation foundation model trained on 1.57M image-mask pairs across 10 imaging modalities and 30+ cancer types (Nature Communications 2024)

Idle4.4K1 year ago
Jupyter Notebook
Apache-2.0

General-purpose pathology foundation model pretrained on 100K+ diagnostic whole-slide images across 20 major tissue types, achieving state-of-the-art transfer learning across 30+ clinical tasks and serving as a universal feature extractor for digital pathology (Mahmood Lab, 722+ stars)

Idle7741 year ago
Jupyter Notebook
NOASSERTION

Kolmogorov-Arnold Networks with learnable activation functions on edges instead of fixed node activations, achieving strong performance in function fitting, PDE solving, and scientific discovery with enhanced interpretability as an alternative to MLPs (MIT, 16.3K+ stars, 2024)

Idle16.3K1 year ago
Jupyter Notebook
MIT

Chemical language model

Idle5011 year ago
Jupyter Notebook
MIT

Deep learning-based protein sequence design (inverse folding) from backbone structures, achieving 52.4% sequence recovery vs 32.9% for Rosetta, core tool in modern protein design pipelines (Baker Lab, Science 2022)

Stale1.8K2 years ago
Jupyter Notebook
MIT

Neural differential equations in PyTorch

Stale1.6K2 years ago
Jupyter Notebook
Apache-2.0

Climate data benchmark for ML models

Stale1152 years ago
Jupyter Notebook
MIT

Generative pre-training for genomics

Stale3242 years ago
Jupyter Notebook

First system to make novel, verifiable scientific discoveries by pairing LLMs with evolutionary search, solving open problems in combinatorics (cap set problem) and discovering faster matrix multiplication algorithms

Stale1.1K2 years ago
Jupyter Notebook
Apache-2.0

Weather prediction benchmark

Stale8372 years ago
Jupyter Notebook
MIT

Large language model for science

Stale2.7K3 years ago
Jupyter Notebook
Apache-2.0