Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

176 of 6,565 resources

Showing 51100

Comprehensive Claude Code skill suite covering the full academic pipeline from deep research and paper writing to multi-perspective peer review, revision, and finalization; features multi-agent teams, PRISMA systematic review, style calibration, claim-level citation audits, integrity gates, and human-in-the-loop safeguards (38K+ stars, CC BY-NC 4.0, 2026)

Active38.4K1 month ago
Python
NOASSERTION
Active81 month ago
Makefile
NOASSERTION

Flow-based generative model for atomistic protein binder design with test-time optimization, SOTA on binder benchmarks (ICLR 2026 Oral, NVIDIA)

Active3991 month ago
Python
NOASSERTION

Latent-space probabilistic denoising diffusion model for predicting coarse-grained conformational ensembles of intrinsically disordered proteins and regions from sequence, with GPU/CPU inference, trajectory export, and FAISS-based similarity search (67+ stars, LGPL-3.0)

Active731 month ago
Jupyter Notebook
NOASSERTION

JavaScript genome browser that is highly customizable via plugins and track customizations.

Active4741 month ago
JavaScript
NOASSERTION

First scientific ML benchmark with paired real-world measurements and matched numerical simulations for complex physical systems, featuring 5 scenarios, 700+ trajectories, 10 baseline models, and 9 evaluation metrics with HuggingFace datasets and model checkpoints (Westlake University, CC BY-NC 4.0)

Active1121 month ago
Python
NOASSERTION

A molecule manipulation library.

Active2371 month ago
Python
NOASSERTION

RFdiffusion is an open source method for structure generation, with or without conditional information (a motif, target etc).

Active3K1 month ago
Python
NOASSERTION

A compressor of common genomic file formats (BAM, CRAM, FASTQ, VCF etc).

Active1851 month ago
C
NOASSERTION

the wavefront alignment algorithm (WFA) which expoit sequence similarity to speed up alignment

Active2241 month ago
C
NOASSERTION

atomate2 is a library of computational materials science workflows.

Active3291 month ago
Python
NOASSERTION

The Data Privacy Vocabulary provides an ontology (classes and properties) and taxonomies of concepts to represent information regarding how personal data is processed in the form of an ontology or a knowledge graph.

Active791 month ago
HTML
NOASSERTION

Oxford Nanopore's official deep-learning basecaller for nanopore sequencing, converting raw electrical signals into DNA/RNA sequences with integrated modified-base (methylation) detection and efficient CPU/GPU inference; foundational tool for long-read genomics, epigenetics, and real-time sequencing analysis (nanoporetech, 846+ stars, actively maintained)

Active8471 month ago
C++
NOASSERTION

Physics-informed neural networks in Julia

Active1.2K1 month ago
Julia
NOASSERTION

Ontologies that aim to provide semantic specifications for units of measure, quantity kind, dimensions and data types.

Active1581 month ago
HTML
NOASSERTION

The modern C++ library for sequence analysis.

Active4581 month ago
C++
NOASSERTION

SOTA multimodal document parsing with 1.2B parameters outperforming GPT-4o, converts PDFs to LLM-ready Markdown/JSON

Active74.3K1 month ago
Python
NOASSERTION

An issue on the UBERON GitHub Issue tracker

Active1581 month ago
Emacs Lisp
NOASSERTION

Diffusion MR imaging.

Active8311 month ago
Python
NOASSERTION

Simulations of spiking neural networks.

Active1.2K1 month ago
Python
NOASSERTION

Toolkit for large-scale whole-slide image processing supporting 22+ patch encoders (UNI, CONCH, Virchow, H-Optimus-0, etc.), slide encoders (TITAN, GigaPath, PRISM, CHIEF, Madeleine, Feather), tissue segmentation, and multi-GPU inference with end-to-end pipeline and smart resume for standardized deployment of computational pathology foundation models (Mahmood Lab, Harvard Medical School, 553+ stars)

Active6001 month ago
Python
NOASSERTION

98B-parameter frontier generative model jointly reasoning over protein sequence, structure, and function, trained on 2.78 billion proteins; generated a novel fluorescent protein (esmGFP) with only 58% sequence identity to known GFPs (EvolutionaryScale, 2024)

Active2.8K1 month ago
Jupyter Notebook
NOASSERTION

Open software framework for Engineering AI built on transformer building blocks, enabling teams to build, train, and operate industrial simulation models across engineering verticals; includes ready-to-use recipes for CFD (AB-UPT on DrivAerML), external aerodynamics, and heat transfer (234+ stars, ENPL non-commercial license, 2026)

Active2341 month ago
Python
NOASSERTION

All-atom generative world model for all-to-all biomolecular interaction design, enabling cross-modality generation of proteins, nucleic acids, small molecules, and cyclic peptides with fine-grained epitope-level control and 2-4 orders of magnitude faster design throughput than modality-specific baselines (316+ stars, Apache 2.0)

Active3271 month ago
Python
NOASSERTION

It is a web-application for visual and interactive gene expression analysis. Phantasus is based on Morpheus – a web-based software for heatmap visualisation and analysis, which was integrated with an R environment via OpenCPU API. Aside from basic visualization and filtering methods, R-based methods such as k-means clustering, principal component analysis or differential expression analysis with limma package are supported.

Active451 month ago
HTML
NOASSERTION

Learnable latent embeddings for joint behavioral and neural analysis, enabling consistent and interpretable mapping of neural activity to behavior across modalities, species, and experiments (EPFL & Harvard, 1K+ stars)

Active1.1K1 month ago
Python
NOASSERTION

Vision foundation model for the tree of life, pretrained on diverse biological imagery across taxa for zero-shot species identification, trait extraction, and biodiversity research (Ohio State University Imageomics Institute)

Active2671 month ago
Python
NOASSERTION

Turn any AI agent into a life science expert with NVIDIA BioNeMo skills, enabling agentic workflows for drug discovery, protein engineering, and biomolecular design (329+ stars, Apache 2.0 / CC-BY-4.0, 2026)

Active3312 months ago
Python
NOASSERTION

Meta FAIR's foundation model of vision, audition, and language for in-silico neuroscience, predicting fMRI brain responses to naturalistic multimodal stimuli (video, audio, text) through unified Transformer architecture mapped to the cortical surface (2026)

Active3.1K2 months ago
Jupyter Notebook
NOASSERTION

Foundation model for universal prompt-driven medical image segmentation extending SAM3 to clinical imaging, supporting 2D public benchmarks and 3D training/evaluation with text and box prompts; pretrained weights available on HuggingFace (189+ stars)

Active1902 months ago
Python
NOASSERTION

DANTE is a software tool for genotyping and characterizing tandem repeats (TRs) from both second- and third-generation sequencing data. It supports the analysis of short-read massively parallel sequencing (sr-MPS) and long-read massively parallel sequencing (lr-MPS), enabling accurate repeat characterization across a wide range of loci. A key feature of DANTE is its ability to determine genotypes at nucleotide resolution, including the characterization and phasing of complex repeat motifs. For sr-MPS data, the tool determines allele size and sequence composition of alleles for which spanning reads are generated. In addition, it identifies alleles that exceed the sequencing read length by estimating their presence from partial read evidence and supports the visualisation of the sequence composition of partial reads. For lr-MPS data, where complete repeat regions are typically sequenced, DANTE determines the allele size and sequence composition of identified alleles.

Active12 months ago
Rust
NOASSERTION

SPAdes (St. Petersburg genome assembler) is an assembly toolkit containing various assembly pipelines and the de-facto standard for prokaryotic genome assemblies.

Active9552 months ago
C++
NOASSERTION

Access to Biological Web Services from Python.

Active3402 months ago
Python
NOASSERTION

Ontology, part of the SI Reference Point, covering measurement units (SI base units and SI units with special names) and prefixes.

Active182 months ago
NOASSERTION

Sparse identification of nonlinear dynamics

Active1.9K2 months ago
Python
NOASSERTION

Closed-loop multi-agent system from hypothesis to verification across 12 scientific tasks, #1 on MLE-Bench (36.44%)

Active1.4K2 months ago
Python
NOASSERTION

A swiss army knife for manipulating and editing PDB files.

Active4562 months ago
Python
NOASSERTION

Lifecycle-Aware Memory (LAM) primitive and benchmark for long-horizon research agents, achieving 65.95% mean reproduction on PaperBench and 94.66% on SurveyBench through Capital Chunk Memory (CCM) with versioned content, structural multi-hop relevance, and provenance-grounded composition; 10 peer-reviewed acceptances at FSE/ICML/TOSEM/AEI/ICoGB (1.3K+ stars)

Active1.3K2 months ago
TeX
NOASSERTION

This package provides a periodic table of the elements with support for mass, density and xray/neutron scattering information.

Active1732 months ago
Python
NOASSERTION

A small language for defining pipeline stages and linking them together to make pipelines.

Active2422 months ago
Groovy
NOASSERTION

Simulation of large-scale brain models

Active9353 months ago
Python
NOASSERTION

Biological vision foundation model trained on TreeOfLife-200M, yielding extraordinary accuracy on diverse biological visual tasks including habitat classification and trait prediction despite a narrow training objective (Ohio State University Imageomics Institute)

Active763 months ago
Python
NOASSERTION

Minimap2 is an pairwise aligner for genomic and spliced nucleotide sequences. It can perform the assembly-to-assembly alignment, and works with gzip'd FASTQ, FASTA formats. It also finds overlaps between long-reads.

Active2.2K3 months ago
C
NOASSERTION

All-atom biomolecular structure prediction for protein-nucleic acid-small molecule-metal ion complexes, enabling accurate modeling of covalent modifications and assemblies beyond proteins (Baker Lab, Science 2024)

Active8143 months ago
Python
NOASSERTION
Active183 months ago
NOASSERTION

Provides methods to convert between Python AnnData objects and SingleCellExperiment objects. These are primarily intended for use by downstream Bioconductor packages that wrap Python methods for single-cell data analysis. It also includes functions to read and write H5AD files used for saving AnnData objects to disk.

Active2093 months ago
R
NOASSERTION

Benchmark quantifying end-to-end autonomous AI research abilities of LLM agents across 20 tasks from SOTA machine learning papers spanning NLP, code, math, biochemical modelling, and time series forecasting, with normalized score metrics against human SOTA and HuggingFace dataset

Active1033 months ago
Python
NOASSERTION

First physics-aligned interactive benchmark for LLM agents in engineering construction, designing rockets/cars/bridges in physics simulator with 3D spatial geometry library

Active963 months ago
Python
NOASSERTION

Arc Institute's single-cell foundation model enabling in-context learning at inference time via a novel tabular attention architecture, trained on 150M uniformly-preprocessed cells for generalizing biological effects and generating unseen cell profiles in novel contexts (2025)

Active1423 months ago
Jupyter Notebook
NOASSERTION

Eukaryotic Genome Annotation Pipeline-External caller scripts and documentation

Active2063 months ago
Nextflow
NOASSERTION