Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

160 of 7,064 resources

Showing 101–150

First physics-aligned interactive benchmark for LLM agents in engineering construction, designing rockets/cars/bridges in physics simulator with 3D spatial geometry library

Active985 months ago
Python
NOASSERTION

Arc Institute's single-cell foundation model enabling in-context learning at inference time via a novel tabular attention architecture, trained on 150M uniformly-preprocessed cells for generalizing biological effects and generating unseen cell profiles in novel contexts (2025)

Active1625 months ago
Jupyter Notebook
NOASSERTION

Benchmark evaluating AI agents on 75 curated Kaggle-style ML engineering competitions with reproducible Docker-based grading harness, human baselines, and end-to-end task lifecycle, used as a primary benchmark for autonomous ML research agents (e.g., InternAgent #1 at 36.44%)

Active1.7K5 months ago
Python
NOASSERTION

toscca is an R package to perform Thresholded Ordered Sparse Canonical Correlation Analysis (TOSCCA).

Active15 months ago
R
NOASSERTION

Dataset and benchmarking framework integrating histology and spatial transcriptomics, enabling multimodal analysis of whole-slide images with matched spatial gene expression for advancing computational pathology and tissue microenvironment research (Mahmood Lab, Harvard Medical School, 411+ stars)

Active4355 months ago
Jupyter Notebook
NOASSERTION

Unified latent diffusion transformer that jointly generates periodic crystals and non-periodic molecules, scaling to 500M parameters with SOTA results on QM9, MP20, and GEOM-DRUGS (Meta FAIR, ICML 2025, 310+ stars)

Active3185 months ago
Python
NOASSERTION

De novo assembler for single molecule sequencing reads using repeat graphs.

Idle9516 months ago
C
NOASSERTION

Baidu's open-source reproduction of AlphaFold3 in PaddlePaddle, providing pretrained weights and inference pipelines for unified biomolecular structure prediction across proteins, nucleic acids, ligands, ions, and post-translational modifications within the PaddleHelix biocomputing platform (Baidu, bioRxiv 2024)

Idle1.1K6 months ago
Python
NOASSERTION

Hybrid deep learning and alignment-based tool for identifying viruses, plasmids, and other mobile genetic elements in isolates, metagenomes, and metatranscriptomes, combining neural-network gene-content classifiers with nucleotide-sequence signatures; also performs viral taxonomic assignment, provirus detection in host genomes, and functional annotation, with precomputed databases of 200K+ viral and 1M+ plasmid genomes and web apps on Galaxy and NMDC EDGE (Berkeley Lab & DOE Joint Genome Institute, 334+ stars, actively maintained)

Idle3346 months ago
Python
NOASSERTION

Genetic variant annotation and effect prediction toolbox.

Idle3137 months ago
Java
NOASSERTION

Multimodal deep learning framework integrating peptide-MHC protein sequence, structure, and biochemical properties to predict class-I immunogenicity for infectious disease epitopes and cancer neoepitopes with cancer-wildtype contrastive learning, enabling personalized vaccine design (Krishnaswamy Lab, Yale University)

Idle487 months ago
Python
NOASSERTION

Foundation models for genomics and transcriptomics pretrained on 3,000+ human genomes and 850+ diverse species, enabling chromatin accessibility prediction, splice site detection, and promoter classification across multiple model scales (InstaDeep, NVIDIA & TUM, Nature Methods 2023)

Idle9197 months ago
Jupyter Notebook
NOASSERTION

Euclidean neural networks for arbitrary point transformations enabling E(3)-equivariant deep learning, foundational library for building geometry-aware neural networks in molecular dynamics, materials science, and physics

Idle1.3K7 months ago
Python
NOASSERTION

Self-supervised vision foundation model for generalized structural brain MRI analysis, pretrained on ~49,000 scans from diverse datasets and generalizing across brain age prediction, dementia/MCI classification, IDH mutation detection, glioma survival prediction, time-to-stroke estimation, MR sequence classification, and brain tumor segmentation; outperforms task-specific models especially with limited training data (Mass General Brigham & Harvard Medical School, 129+ stars)

Idle1578 months ago
Python
NOASSERTION

Lightweight supervised slide foundation model with 0.9M parameters pretrained on 24K whole-slide images for pan-cancer morphological classification, achieving competitive performance with much larger self-supervised models (TITAN, GigaPath) while enabling finetuning on consumer-grade GPUs; includes standardized MIL implementations and benchmarking across 15+ classification tasks (Mahmood Lab, Harvard Medical School, 153+ stars)

Idle1588 months ago
Python
NOASSERTION

scQTLtools is a comprehensive R/Bioconductor package that facilitates end-to-end single-cell eQTL analysis, from preprocessing to visualization

Idle78 months ago
R
NOASSERTION

A [Jupyter](https://jupyter.org/) widget to interactively view molecular structures and trajectories.

Idle9278 months ago
Jupyter Notebook
NOASSERTION

A new clustering algorithm, "binary cut", for clustering similarity matrices of functional terms is implemeted in this package. It also provides functions for visualizing, summarizing and comparing the clusterings.

Idle1288 months ago
R
NOASSERTION

SCENIC+ is a python package to build gene regulatory networks (GRNs) using combined or separate single-cell gene expression (scRNA-seq) and single-cell chromatin accessibility (scATAC-seq) data.

Idle2638 months ago
Jupyter Notebook
NOASSERTION

Official implementation of the second-generation fully autonomous scientific discovery system, extending the original with agentic tree search and reduced template dependency to achieve workshop-level accepted papers (6.7K+ stars, 2025)

Idle7.3K9 months ago
Python
NOASSERTION

First fully autonomous open-ended scientific discovery system with official implementation: hypothesis→experiment→writing→review simulation (13.8K+ stars, 2024)

Idle14.5K9 months ago
Jupyter Notebook
NOASSERTION

This package provides functionality to run a number of tasks in the differential expression analysis workflow. This encompasses the most widely used steps, from running various enrichment analysis tools with a unified interface to creating plots and beautifying table components linking to external websites and databases. This streamlines the generation of comprehensive analysis reports.

Idle010 months ago
R
NOASSERTION

Biocaml aims to be a high-performance user-friendly library for Bioinformatics.

Idle12310 months ago
OCaml
NOASSERTION

Graph neural network operating entirely at the atomic level for protein-ligand conformational ensemble prediction and docking, generating diverse solutions through rapid stochastic denoising to model conformational heterogeneity (Baker Lab, bioRxiv 2025)

Idle26711 months ago
Python
NOASSERTION

Conversational data analysis using natural language

Idle23.8K11 months ago
Python
NOASSERTION

Structural variant and indel caller for mapped sequencing data.

Archived46712 months ago
C++
NOASSERTION

AI-assisted mutation nomination approach optimizing protein function by integrating structural and evolutionary constraints into protein inverse folding models, compatible with ProteinMPNN, LigandMPNN, ESM-IF1, and SaProt (Chinese Academy of Sciences, 359+ stars)

Idle1.1K1 year ago
Jupyter Notebook
NOASSERTION

Another list focuses on Python stuff related to Chemistry.

Idle1.4K1 year ago
NOASSERTION

Cheminformatic extension for the SQLAlchemy database.

Idle401 year ago
Python
NOASSERTION

SIMD C library for global, semi-global, and local pairwise sequence alignments

Idle2881 year ago
C
NOASSERTION

NIST's open-source platform for data-driven atomistic materials design, integrating DFT datasets (JARVIS-DFT), machine learning property prediction (JARVIS-ML), and a comprehensive leaderboard for benchmarking materials AI methods across the periodic table (384+ stars)

Idle4011 year ago
Python
NOASSERTION

Large-scale flow-based protein backbone generator utilizing hierarchical fold class labels for conditioning with a tailored scalable transformer architecture, enabling controllable de novo protein design (264+ stars)

Idle2761 year ago
Python
NOASSERTION

DeepSeek's open-source large language model for formal theorem proving in Lean 4, integrating informal and formal mathematical reasoning through recursive subgoal decomposition and reinforcement learning powered by DeepSeek-V3, with open weights and ProverBench evaluation (2025)

Idle1.3K1 year ago
NOASSERTION

In silico directed evolution framework using few-shot active learning to optimize protein activities, enabling rapid protein engineering with minimal experimental data (352+ stars, 2023)

Idle3761 year ago
Python
NOASSERTION

Large language-and-vision assistant for biomedicine, instruction-tuned on GPT-4-generated biomedical multimodal instruction-following data to enable conversational visual question answering over radiology, pathology, and microscopy images, establishing open recipes for adapting general vision-language models to the biomedical domain (Microsoft Research & University of Washington, 2.2K+ stars)

Idle2.2K1 year ago
Python
NOASSERTION

GRIDSS: the Genomic Rearrangement IDentification Software Suite.

Idle2861 year ago
Java
NOASSERTION

Systematic medical RAG toolkit for question answering over PubMed, StatPearls, textbooks, and Wikipedia, supporting multiple retrievers, domain LLMs, and follow-up-query workflows for benchmarked clinical/biomedical QA (ACL Findings 2024)

Idle6001 year ago
Python
NOASSERTION

General-purpose pathology foundation model pretrained on 100K+ diagnostic whole-slide images across 20 major tissue types, achieving state-of-the-art transfer learning across 30+ clinical tasks and serving as a universal feature extractor for digital pathology (Mahmood Lab, 722+ stars)

Idle7741 year ago
Jupyter Notebook
NOASSERTION

Vision-language pathology foundation model using contrastive learning on histopathology image-text pairs, enabling zero-shot classification, slide-level retrieval, and multimodal reasoning across diverse cancer types (Mahmood Lab, 494+ stars)

Idle5341 year ago
Python
NOASSERTION

A database system designed to store, organize, and manage large-scale nucleotide sequencing read data (like PacBio reads) for the Dazzler genome assembler

Idle361 year ago
C
NOASSERTION

SKESA is a de-novo sequence read assembler for microbial genomes. It uses conservative heuristics and is designed to create breaks at repeat regions in the genome. This leads to excellent sequence quality without significantly compromising contiguity.

Idle1271 year ago
C++
NOASSERTION

Universal chart comprehension and reasoning model

Stale1362 years ago
Python
NOASSERTION

Functions to summarize DNA methylation data using regional principal components. Regional principal components are computed using principal components analysis within genomic regions to summarize the variability in methylation levels across CpGs. The number of principal components is chosen using either the Marcenko-Pasteur or Gavish-Donoho method to identify relevant signal in the data.

Stale42 years ago
R
NOASSERTION

[RDKit](http://www.rdkit.org/) and [OSRA](https://cactus.nci.nih.gov/osra/) in the [Bottle](http://bottlepy.org/docs/dev/) on [Tornado](http://www.tornadoweb.org/en/stable/).

Archived502 years ago
Python
NOASSERTION

Circlator is a tool to circularize genome assemblies. It will attempt to identify each circular sequence and output a linearised version of it. It does this by assembling all reads that map to contig ends and comparing the resulting contigs with the input assembly.

Stale2592 years ago
Python
NOASSERTION

k-mer counting, filtering, and graph traversal.

Stale7892 years ago
Python
NOASSERTION

NOVOPlasty - The organelle assembler and heteroplasmy caller. NOVOPlasty is a de novo assembler and heteroplasmy/variance caller for short circular genomes..

Stale2012 years ago
Perl
NOASSERTION

A VCF Parser for Python.

Stale4193 years ago
Python
NOASSERTION

Displaying sequence statistics for next-generation sequencing.

Stale253 years ago
C
NOASSERTION

Educational resource on performing RNA-seq analysis in the cloud using Amazon AWS cloud services. Topics include preparing the data, preprocessing, differential expression, isoform discovery, data visualization, and interpretation.

Stale1.4K3 years ago
R
NOASSERTION