Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

2,030 of 6,590 resources

Showing 51100

Interactive and hardware-agnostic SDK for laboratory automation, enabling programmatic control of liquid handlers, plate readers, and other lab instruments across multiple vendors; foundational infrastructure for self-driving laboratories and AI-driven experimental execution (447+ stars)

Active5031 week ago
Python
MIT

Provides functionality for producing geometric representations of protein and RNA structures, and biological interaction networks.

Active1.2K1 week ago
Jupyter Notebook
MIT

spoQC is a modular framework for multimodal quality control (QC) of imaging-based spatially resolved transcriptomics (SRT). It independently evaluates cell segmentation, imaging, and transcript data to identify high-quality regions (HQRs) across entire tissue sections. In addition, spoQC uses Markov random fields (MRFs) to incorporate spatial dependencies and generate spatially refined QC masks.

Active01 week ago
Python
MIT

The NCBI Gene Expression Omnibus (GEO) is a public repository of microarray data. Given the rich and varied nature of this resource, it is only natural to want to apply BioConductor tools to these data. GEOquery is the bridge between GEO and BioConductor.

Active1161 week ago
R
MIT

Google DeepMind's unified DNA sequence foundation model predicting molecular consequences of genetic variants from single-base resolution up to 1 megabase context, jointly outputting thousands of regulatory tracks (RNA expression, splicing, chromatin accessibility, TF binding, contact maps) for human and mouse genomes via a Python client and non-commercial API (2025)

Active2K1 week ago
Python
Apache-2.0

A library and command-line tool for building and analyzing complex homogeneous microkinetic models from quantum chemistry calculations, with support for quasi-harmonic thermochemistry, quantum tunnelling corrections, molecular symmetries and more.

Active641 week ago
Python
MIT

R package for analysis of transcript and translation features through manipulation of sequence data and NGS data like Ribo-Seq, RNA-Seq, TCP-Seq and CAGE. It is generalized in the sense that any transcript region can be analysed, as the name hints to it was made with investigation of ribosomal patterns over Open Reading Frames (ORFs) as it's primary use case. ORFik is extremely fast through use of C++, data.table and GenomicRanges. Package allows to reassign starts of the transcripts with the use of CAGE-Seq data, automatic shifting of RiboSeq reads, finding of Open Reading Frames for whole genomes and much more.

Active381 week ago
R
MIT

Julia differential equations suite

Active3.1K1 week ago
Julia
NOASSERTION

This package is a wrapper of Integrative Genomics Viewer (IGV). It comprises an htmlwidget version of IGV. It can be used as a module in Shiny apps.

Active431 week ago
R
MIT

An object-oriented, webGL based JavaScript library for online molecular visualization.

Active1K1 week ago
Jupyter Notebook
NOASSERTION

A small language for defining pipeline stages and linking them together to make pipelines.

Active2421 week ago
Groovy
NOASSERTION

A package for working with nuclear magnetic resonance (NMR) data including functions for reading common binary file formats and processing NMR data.

Active2721 week ago
Python
BSD-3-Clause

OpenTFRaw is a standalone, cross-platform reader for Thermo Fisher Scientific .raw mass-spectrometry files, implemented in pure Rust with no dependency on vendor DLLs or .NET. Python bindings built on PyO3 return NumPy arrays for spectral data, straightforward to load into Pandas or Polars. Covers format versions 8 through 66 (LCQ Classic through Orbitrap Astral and modern TSQ instruments), supporting both centroid and profile spectra.

Active131 week ago
Rust
NOASSERTION

OpenWRaw is a standalone, cross-platform reader for Waters MassLynx .raw acquisition directories, implemented in pure Rust with no dependency on vendor DLLs. Python bindings built on PyO3 expose functions, scans, and ion-mobility data as native Python objects from Waters QTof and SYNAPT instrument families, ready to be assembled into a Pandas or Polars DataFrame.

Active31 week ago
Rust
NOASSERTION

OpenTimsTDF is a standalone, cross-platform reader for Bruker timsTOF .tdf and .tdf_bin acquisition files, implemented in pure Rust with no dependency on vendor SDKs. Python bindings built on PyO3 expose frame, scan, and peak data as native Python objects, providing ion-mobility-aware access that can be assembled into a Pandas or Polars DataFrame.

Active21 week ago
Rust
NOASSERTION

Toolkit for large-scale whole-slide image processing supporting 22+ patch encoders (UNI, CONCH, Virchow, H-Optimus-0, etc.), slide encoders (TITAN, GigaPath, PRISM, CHIEF, Madeleine, Feather), tissue segmentation, and multi-GPU inference with end-to-end pipeline and smart resume for standardized deployment of computational pathology foundation models (Mahmood Lab, Harvard Medical School, 553+ stars)

Active6241 week ago
Python
NOASSERTION

Open-source, all-atom biomolecular foundation model that turns co-folding into a scalable engine for structure prediction, design, and optimization across proteins, nucleic acids, and small molecules in drug discovery; ranked first on PXMeter-AB, FoldBench-AB, and 2026ARK-AB antibody-antigen benchmarks (263+ stars, Apache 2.0)

Active4231 week ago
Python
Apache-2.0

Fits second-order autoregressive AR(2) models to gene expression time series and reports the eigenvalue modulus |lambda|, a single statistic quantifying temporal persistence: how strongly a gene's recent past constrains its next value. Ranks genes into a clock/target/background hierarchy and reports correlation length, half-life and root type (real or complex) per gene.

Active01 week ago
Python
Other

E(3)-equivariant neural network interatomic potentials achieving DFT accuracy with up to 1000× less training data than invariant models, foundational architecture behind MACE and Allegro (Harvard, MIT, Nature Communications 2022)

Active9541 week ago
Python
MIT

First unified Lean 4 framework for neural-network specification, execution, and verification; tensor shapes are part of the types, models are executable Lean programs, and the same definitions can be used by training code, graph transformations, certificate checkers, and proofs with CPU/CUDA backends (123+ stars, MIT License)

Active1241 week ago
Lean
MIT

A quantum chemistry package written in Python.

Active791 week ago
Python
Apache-2.0

Sequence manipulation toolkit for FASTA/FASTQ files written in Nim.

Active1291 week ago
Nim
GPL-3.0

a specification for describing analysis workflows and tools that are portable and scalable across a variety of software and hardware environments, from workstations to cluster, cloud, and high performance computing (HPC) environments.

Active1.5K1 week ago
Common Workflow Language
Apache-2.0

Collection of SKILLS.md guiding AI coding agents (Claude Code, OpenAI Codex, Google Gemini, OpenCode, OpenClaw) through common bioinformatics workflows from basic sequence manipulation to advanced analyses such as single-cell RNA-seq and population genetics; evaluated on the Bio-Task Bench dataset (GPTomics, 969+ stars, MIT License, 2026)

Archived1.2K1 week ago
Python
MIT

Co-create PowerPoint presentations with Generative AI from documents or topics

Active3711 week ago
Python
MIT

scider is an user-friendly R package providing functions to model the global density of cells in a slide of spatial transcriptomics data. All functions in the package are built based on the SpatialExperiment object, allowing integration into various spatial transcriptomics-related packages from Bioconductor. After modelling density, the package allows for several downstream analysis, including colocalization analysis, boundary detection analysis and differential density analysis.

Active121 week ago
R
GPL-3.0

Exact, validated excision of coordinate-defined genomic regions from transposed NEXUS matrices.

Active11 week ago
Python
MIT

Semi-autonomous AI scientist for scientific theory discovery and verifiable goal solving, using adversarial review-refinement loops and evolution-inspired candidate populations; integrates with Claude Code, Gemini CLI, Antigravity, and Codex harnesses (Imbue, 31+ stars, AGPL-3.0, 2026)

Active321 week ago
Python
AGPL-3.0

Provides utilities for identifying drug-target interactions for sets of small molecule or gene/protein identifiers. The required drug-target interaction information is obained from a local SQLite instance of the ChEMBL database. ChEMBL has been chosen for this purpose, because it provides one of the most comprehensive and best annotatated knowledge resources for drug-target information available in the public domain.

Active21 week ago
R
Artistic-2.0

Hand-curated Snakemake pipelines to combine identifier cross-references from multiple sources across dozens of biomedical types, including anatomical entities, diseases and phenotypes, genes and proteins and many others.

Active171 week ago
Python
MIT

BRANCHSNV reports strict clade-exclusive nucleotide markers separately from single-nucleotide substitutions reconstructed on a selected edge of a rooted phylogenetic tree, while retaining ambiguity across equally parsimonious ancestral-state reconstructions.

Active11 week ago
Python
MIT

Multi-type data labeling and annotation tool

Active28.1K1 week ago
TypeScript
Apache-2.0

Python toolkit for fine-tuning geospatial foundation models

Active8441 week ago
Python
Apache-2.0

High-Throughput Molecular Dynamics: Programming Environment for Molecular Discovery.

Active2761 week ago
Rich Text Format
NOASSERTION

Low- and high-level wrappers for Gemma's RESTful API. They enable access to curated expression and differential expression data from over 10,000 published studies. Gemma is a web site, database and a set of tools for the meta-analysis, re-use and sharing of genomics data, currently primarily targeted at the analysis of gene expression profiles.

Active101 week ago
R
Apache-2.0+

First bioinformatics-native AI agent skill library enabling local-first, reproducible genomic and population-genetics research workflows built on OpenClaw (871+ stars, MIT License, 2026)

Active1.1K1 week ago
Python
NOASSERTION

Deterministic, rule-based variant interpretation platform for clinical genetics laboratories. Automates ACMG/AMP 2015 classification using a Bayesian point-based framework (Tavtigian et al. 2018) with BayesDel ClinGen SVI-calibrated thresholds (Pejaver et al. 2022). Integrates 8 reference databases (gnomAD v4.1, ClinVar, dbNSFP 4.9c, SpliceAI, gnomAD Constraint, HPO, ClinGen, Ensembl VEP). Analyzes nuclear and mtDNA variants, structural and copy-number variants (SV/CNV), with trio/family and cohort analysis. Supports HPO-based phenotype matching, biomedical literature mining across 2M+ PubMed publications, and structured clinical report generation. AI assists in evidence synthesis but does not make classification decisions. EU-hosted on dedicated infrastructure in Helsinki, Finland (GDPR-compliant).

Active01 week ago
Python
Proprietary

Genomic data analyses requires integrated visualization of known genomic information and new experimental data. Gviz uses the biomaRt and the rtracklayer packages to perform live annotation queries to Ensembl and UCSC and translates this to e.g. gene/transcript structures in viewports of the grid graphics package. This results in genomic information plotted together with your data.

Active921 week ago
R
Artistic-2.0

A Python package for protein dynamics analysis

Active5542 weeks ago
Python
NOASSERTION

Modular toolchain for an extensible and customizable ETL pipeline that extracts, transforms, and loads clinical data and medical imaging metadata, applying dataset-specific mappings to generate outputs compatible with the EUCAIM Common Data Model (CDM). Its design aims to minimize manual data preparation efforts and facilitate customization and integration with other components, such as data quality assurance tools. Containerized, currently supports input datasets in CSV, JSON, XLSX.

Active12 weeks ago
Python

PanAbyss is a tool for exploring and visualizing pangenome graphs. It allows users to search for and display regions of a pangenome using coordinates on a reference individual or based on annotations. It also enables searching for regions associated with a selected set of individuals (for example, those linked to a phenotype), computing proximity trees, and retrieving sequences from a given region.

Active32 weeks ago
GPL-3.0

edfcore is a zero-dependency TypeScript library for reading EDF, EDF+, BDF, and BDF+ physiological recordings in browser and Node.js applications. It provides programmatic access to biosignal samples, channel metadata, per-channel sampling rates, physical units, annotations, and discontinuous recording timelines.

Active12 weeks ago
JavaScript
MIT

A collection of object-oriented software tools for problems involving chemical kinetics, thermodynamics, and transport processes.

Active8382 weeks ago
C++
NOASSERTION

Open-source image analysis toolkit for high-throughput plant phenotyping, extracting morphological, color, and texture traits from RGB, hyperspectral, and thermal imagery with modular Python workflows for crop improvement, stress detection, and plant biology research (Donald Danforth Plant Science Center, 795+ stars, MPL-2.0)

Active8142 weeks ago
Python
MPL-2.0

A transparent, unit-aware calculator for the mathematical relationship between peptide mass, target concentration and solution volume. It normalizes mg, micrograms, mL and microlitres, shows the formula and includes a reference syringe visualization. Research-use-only software: it does not select a solvent, validate a laboratory method, calculate a dose or provide administration guidance.

Active02 weeks ago
MIT

A benchmark for ML-guided high-throughput materials discovery.

Active2462 weeks ago
Python
MIT

Unified Python framework for bulk, single-cell, and spatial RNA-seq multi-omics analysis with deep learning deconvolution (VAE) and graph neural networks, bridging Bindea, Bindea, scanpy and squidpy ecosystems (Nature Communications 2024)

Active1.2K2 weeks ago
Python
GPL-3.0

A client to simplify fetching predictions from the Koina web service. Koina is a model repository enabling the remote execution of models. Predictions are generated as a response to HTTP/S requests, the standard protocol used for nearly all web traffic.

Active572 weeks ago
R
Apache-2.0

Continuously updated functional re-annotation of the Mycobacterium tuberculosis complex gene set, anchored on the MTBC0 ancestral genome rather than on a single strain. Serves one record per gene combining Pfam domains, ESMFold structures with Foldseek search, protein language-model features, orthology, curated knowledge, protein association networks and intra-species selection inferred from 145209 sequenced genomes, with dated sources and a graded confidence level for every field. Intended as a successor to Mycobrowser, which is no longer maintained.

Active02 weeks ago
Python
CC-BY-4.0

AI-assisted structural engineering workspace for AEC workflows: natural language to structural model, analysis, code-check, and report (171+ stars, MIT License, 2026)

Active1712 weeks ago
Python
MIT