Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

149 of 6,573 resources

Showing 101149

GenBio AI's software stack for the AI-Driven Digital Organism, supporting adaptation and finetuning of multiscale biological foundation models across DNA, RNA, protein, structure, and single-cell tasks with reproducible CLIs and pretrained model zoo (2025)

Idle1186 months ago
Python
NOASSERTION

Foundation models for genomics and transcriptomics pretrained on 3,000+ human genomes and 850+ diverse species, enabling chromatin accessibility prediction, splice site detection, and promoter classification across multiple model scales (InstaDeep, NVIDIA & TUM, Nature Methods 2023)

Idle9016 months ago
Jupyter Notebook
NOASSERTION

Universal pretrained neural network potential with charge and magnetic moment awareness, trained on 1.5M+ Materials Project inorganic structures for charge-informed molecular dynamics and phase diagram prediction (Berkeley, Nature Machine Intelligence 2023 Cover)

Idle3996 months ago
Python
NOASSERTION

Euclidean neural networks for arbitrary point transformations enabling E(3)-equivariant deep learning, foundational library for building geometry-aware neural networks in molecular dynamics, materials science, and physics

Idle1.3K6 months ago
Python
NOASSERTION

Self-supervised vision foundation model for generalized structural brain MRI analysis, pretrained on ~49,000 scans from diverse datasets and generalizing across brain age prediction, dementia/MCI classification, IDH mutation detection, glioma survival prediction, time-to-stroke estimation, MR sequence classification, and brain tumor segmentation; outperforms task-specific models especially with limited training data (Mass General Brigham & Harvard Medical School, 129+ stars)

Idle1406 months ago
Python
NOASSERTION

Lightweight supervised slide foundation model with 0.9M parameters pretrained on 24K whole-slide images for pan-cancer morphological classification, achieving competitive performance with much larger self-supervised models (TITAN, GigaPath) while enabling finetuning on consumer-grade GPUs; includes standardized MIL implementations and benchmarking across 15+ classification tasks (Mahmood Lab, Harvard Medical School, 153+ stars)

Idle1526 months ago
Python
NOASSERTION

A [Jupyter](https://jupyter.org/) widget to interactively view molecular structures and trajectories.

Idle9266 months ago
Jupyter Notebook
NOASSERTION

SCENIC+ is a python package to build gene regulatory networks (GRNs) using combined or separate single-cell gene expression (scRNA-seq) and single-cell chromatin accessibility (scATAC-seq) data.

Idle2617 months ago
Jupyter Notebook
NOASSERTION

Official implementation of the second-generation fully autonomous scientific discovery system, extending the original with agentic tree search and reduced template dependency to achieve workshop-level accepted papers (6.7K+ stars, 2025)

Idle7K8 months ago
Python
NOASSERTION

First fully autonomous open-ended scientific discovery system with official implementation: hypothesis→experiment→writing→review simulation (13.8K+ stars, 2024)

Idle14.3K8 months ago
Jupyter Notebook
NOASSERTION

Biocaml aims to be a high-performance user-friendly library for Bioinformatics.

Idle1239 months ago
OCaml
NOASSERTION

Graph neural network operating entirely at the atomic level for protein-ligand conformational ensemble prediction and docking, generating diverse solutions through rapid stochastic denoising to model conformational heterogeneity (Baker Lab, bioRxiv 2025)

Idle2609 months ago
Python
NOASSERTION

Conversational data analysis using natural language

Idle23.7K10 months ago
Python
NOASSERTION

Structural variant and indel caller for mapped sequencing data.

Archived46810 months ago
C++
NOASSERTION

AI-assisted mutation nomination approach optimizing protein function by integrating structural and evolutionary constraints into protein inverse folding models, compatible with ProteinMPNN, LigandMPNN, ESM-IF1, and SaProt (Chinese Academy of Sciences, 359+ stars)

Idle1.1K10 months ago
Jupyter Notebook
NOASSERTION

Another list focuses on Python stuff related to Chemistry.

Idle1.4K11 months ago
NOASSERTION

Cheminformatic extension for the SQLAlchemy database.

Idle4011 months ago
Python
NOASSERTION

SIMD C library for global, semi-global, and local pairwise sequence alignments

Idle28512 months ago
C
NOASSERTION

NIST's open-source platform for data-driven atomistic materials design, integrating DFT datasets (JARVIS-DFT), machine learning property prediction (JARVIS-ML), and a comprehensive leaderboard for benchmarking materials AI methods across the periodic table (384+ stars)

Idle3901 year ago
Python
NOASSERTION

Large-scale flow-based protein backbone generator utilizing hierarchical fold class labels for conditioning with a tailored scalable transformer architecture, enabling controllable de novo protein design (264+ stars)

Idle2701 year ago
Python
NOASSERTION

DeepSeek's open-source large language model for formal theorem proving in Lean 4, integrating informal and formal mathematical reasoning through recursive subgoal decomposition and reinforcement learning powered by DeepSeek-V3, with open weights and ProverBench evaluation (2025)

Idle1.3K1 year ago
NOASSERTION

In silico directed evolution framework using few-shot active learning to optimize protein activities, enabling rapid protein engineering with minimal experimental data (352+ stars, 2023)

Idle3641 year ago
Python
NOASSERTION

animalcules is an R package for utilizing up-to-date data analytics, visualization methods, and machine learning models to provide users an easy-to-use interactive microbiome analysis framework. It can be used as a standalone software package or users can explore their data with the accompanying interactive R Shiny application. Traditional microbiome analysis such as alpha/beta diversity and differential abundance analysis are enhanced, while new methods like biomarker identification are introduced by animalcules. Powerful interactive and dynamic figures generated by animalcules enable users to understand their data better and discover new insights.

Idle551 year ago
R
NOASSERTION

GRIDSS: the Genomic Rearrangement IDentification Software Suite.

Idle2851 year ago
Java
NOASSERTION

Systematic medical RAG toolkit for question answering over PubMed, StatPearls, textbooks, and Wikipedia, supporting multiple retrievers, domain LLMs, and follow-up-query workflows for benchmarked clinical/biomedical QA (ACL Findings 2024)

Idle5801 year ago
Python
NOASSERTION

General-purpose pathology foundation model pretrained on 100K+ diagnostic whole-slide images across 20 major tissue types, achieving state-of-the-art transfer learning across 30+ clinical tasks and serving as a universal feature extractor for digital pathology (Mahmood Lab, 722+ stars)

Idle7611 year ago
Jupyter Notebook
NOASSERTION

Vision-language pathology foundation model using contrastive learning on histopathology image-text pairs, enabling zero-shot classification, slide-level retrieval, and multimodal reasoning across diverse cancer types (Mahmood Lab, 494+ stars)

Idle5211 year ago
Python
NOASSERTION

Python wrapper for [bedtools](https://github.com/arq5x/bedtools).

Idle3301 year ago
Python
NOASSERTION

A database system designed to store, organize, and manage large-scale nucleotide sequencing read data (like PacBio reads) for the Dazzler genome assembler

Idle361 year ago
C
NOASSERTION

Descriptor computation(chemistry) and (optional) storage for machine learning.

Idle2801 year ago
Python
NOASSERTION

SKESA is a de-novo sequence read assembler for microbial genomes. It uses conservative heuristics and is designed to create breaks at repeat regions in the genome. This leads to excellent sequence quality without significantly compromising contiguity.

Idle1261 year ago
C++
NOASSERTION

This packages simulates spatial transcriptomics data with the mean- variance relationship using a Gaussian Process model per gene.

Idle01 year ago
R
NOASSERTION

Universal chart comprehension and reasoning model

Idle1351 year ago
Python
NOASSERTION

[RDKit](http://www.rdkit.org/) and [OSRA](https://cactus.nci.nih.gov/osra/) in the [Bottle](http://bottlepy.org/docs/dev/) on [Tornado](http://www.tornadoweb.org/en/stable/).

Archived502 years ago
Python
NOASSERTION

Circlator is a tool to circularize genome assemblies. It will attempt to identify each circular sequence and output a linearised version of it. It does this by assembling all reads that map to contig ends and comparing the resulting contigs with the input assembly.

Stale2572 years ago
Python
NOASSERTION

k-mer counting, filtering, and graph traversal.

Stale7912 years ago
Python
NOASSERTION

NOVOPlasty - The organelle assembler and heteroplasmy caller. NOVOPlasty is a de novo assembler and heteroplasmy/variance caller for short circular genomes..

Stale1982 years ago
Perl
NOASSERTION

A VCF Parser for Python.

Stale4192 years ago
Python
NOASSERTION

Displaying sequence statistics for next-generation sequencing.

Stale243 years ago
C
NOASSERTION

Educational resource on performing RNA-seq analysis in the cloud using Amazon AWS cloud services. Topics include preparing the data, preprocessing, differential expression, isoform discovery, data visualization, and interpretation.

Stale1.4K3 years ago
R
NOASSERTION

Open source web framework for small molecule analysis based on Django.

Stale423 years ago
JavaScript
NOASSERTION

Learning nonlinear operators

Stale8314 years ago
Python
NOASSERTION

AI for chemical reaction prediction and synthesis planning

Stale4284 years ago
Python
NOASSERTION

FASTQ/A short-reads pre-processing tools: Demultiplexing, trimming, clipping, quality filtering, and masking utilities.

Stale2024 years ago
C
NOASSERTION

Finds SNP sites from a multi-FASTA alignment file.

Stale2785 years ago
C
NOASSERTION

Customizable pipeline for differential expression analysis with an intuitive GUI.

Stale76 years ago
Java
NOASSERTION

Flexible circular visualization of genome-associated data with BioPerl and SVG.

Stale467 years ago
Perl
NOASSERTION

Horizon chart D3-based JavaScript library for DNA data.

Stale6210 years ago
CSS
NOASSERTION