Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

996 of 6,511 resources

Showing 150

Hand-curated Snakemake pipelines to combine identifier cross-references from multiple sources across dozens of biomedical types, including anatomical entities, diseases and phenotypes, genes and proteins and many others.

Active1720 hours ago
Python
MIT

BRANCHSNV reports strict clade-exclusive nucleotide markers separately from single-nucleotide substitutions reconstructed on a selected edge of a rooted phylogenetic tree, while retaining ambiguity across equally parsimonious ancestral-state reconstructions.

Active11 day ago
Python
MIT

First bioinformatics-native AI agent skill library enabling local-first, reproducible genomic and population-genetics research workflows built on OpenClaw (871+ stars, MIT License, 2026)

Active1.1K1 day ago
Python
NOASSERTION

Deterministic, rule-based variant interpretation platform for clinical genetics laboratories. Automates ACMG/AMP 2015 classification using a Bayesian point-based framework (Tavtigian et al. 2018) with BayesDel ClinGen SVI-calibrated thresholds (Pejaver et al. 2022). Integrates 8 reference databases (gnomAD v4.1, ClinVar, dbNSFP 4.9c, SpliceAI, gnomAD Constraint, HPO, ClinGen, Ensembl VEP). Analyzes nuclear and mtDNA variants, structural and copy-number variants (SV/CNV), with trio/family and cohort analysis. Supports HPO-based phenotype matching, biomedical literature mining across 2M+ PubMed publications, and structured clinical report generation. AI assists in evidence synthesis but does not make classification decisions. EU-hosted on dedicated infrastructure in Helsinki, Finland (GDPR-compliant).

Active01 day ago
Python
Proprietary

PanAbyss is a tool for exploring and visualizing pangenome graphs. It allows users to search for and display regions of a pangenome using coordinates on a reference individual or based on annotations. It also enables searching for regions associated with a selected set of individuals (for example, those linked to a phenotype), computing proximity trees, and retrieving sequences from a given region.

Active32 days ago
Python
NOASSERTION

Open-source image analysis toolkit for high-throughput plant phenotyping, extracting morphological, color, and texture traits from RGB, hyperspectral, and thermal imagery with modular Python workflows for crop improvement, stress detection, and plant biology research (Donald Danforth Plant Science Center, 795+ stars, MPL-2.0)

Active8142 days ago
Python
MPL-2.0

A benchmark for ML-guided high-throughput materials discovery.

Active2462 days ago
Python
MIT

The primary goal of this ontology is to standardize the representation of molecular simulation data, processes, and methodologies across disparate simulation platforms, engines (e.g., GROMACS, AMBER, NAMD), and analysis tools, while ensuring these terms are interoperable with existing life sciences ontologies

Active73 days ago
Python

BondShift: Organic Mechanism Reasoning

Active123 days ago
Python

Continuously updated functional re-annotation of the Mycobacterium tuberculosis complex gene set, anchored on the MTBC0 ancestral genome rather than on a single strain. Serves one record per gene combining Pfam domains, ESMFold structures with Foldseek search, protein language-model features, orthology, curated knowledge, protein association networks and intra-species selection inferred from 145209 sequenced genomes, with dated sources and a graded confidence level for every field. Intended as a successor to Mycobrowser, which is no longer maintained.

Active03 days ago
Python
CC-BY-4.0

AI-assisted structural engineering workspace for AEC workflows: natural language to structural model, analysis, code-check, and report (171+ stars, MIT License, 2026)

Active1713 days ago
Python
MIT

The System Package Data Exchange™ (SPDX®) specification is an open standard designed to represent systems containing software components as Software Bill of Materials (SBOMs). Additionally, SPDX supports AI, data, and security references, making it suitable for a wide range of risk management use cases. This _spdx3_ prefix is for SPDX 3.x versions. For earlier versions, use _spdx.term_.

Active3783 days ago
Python
Community-Spec-1.0

Local Python sequence utilities for nucleotide composition, DNA and RNA reverse complements, NCBI genetic-code translation, six-frame candidate ORF enumeration, and IUPAC motif searches. Computase accepts raw nucleotide strings or one FASTA record and returns structured, bounded results with explicit scientific conventions.

Active04 days ago
Python
MIT

Lightweight Markdown-only skills for autonomous ML research with cross-model review loops, idea discovery, and experiment automation; no framework lock-in, works with Claude Code, Codex, OpenClaw, or any LLM agent (12.8K+ stars, MIT License, 2026)

Active14.7K4 days ago
Python
MIT

Molecular dynamics analysis

Active1.6K4 days ago
Python
NOASSERTION

Local-first, open-source healthcare AI toolkit for clinical NLP and PHI/PII de-identification across 12 languages, running entirely on-device with 1,000+ specialized medical models; provides Python SDK, REST API, Docker deployment, and native Swift apps via OpenMedKit with Apple MLX/CoreML acceleration, supporting HIPAA-aware de-identification with 247 PII checkpoints (3K+ stars, Apache 2.0, arXiv 2508.01630)

Active5K4 days ago
Python
Apache-2.0

Beyond text-to-slides generation with PPTEval multi-dimensional evaluation (EMNLP 2025)

Active4.9K4 days ago
Python
MIT

Probabilistic programming

Active9.7K5 days ago
Python
NOASSERTION

A clinical reasoning assistant for early-stage Alzheimer's assessment. It joins a 3D MRI + biomarker classifier (Vbai-2.6AD) to a reasoning LLM (Gemma 4 12B) inside a single forward pass — the diagnosis is passed as vectors, not text.

Active05 days ago
Python

The information resource registry is a listing of data sources present in the NCATS Data Translator system. Each information resource has an identifier, a short description, and a URL to more information about that resource.

Active65 days ago
Python
Apache-2.0

Human-centered research OS with terminal-first harness and local browser Studio, turning research work into reproducible artifact-backed runs through a 9-stage workflow with human approval gates, resume/rollback controls, and venue-aware manuscript packaging (1K+ stars, 2026)

Active8105 days ago
Python
NOASSERTION

Shared multimodal AI agent layer for geospatial Python packages (leafmap, geoai, geemap, STAC, NASA Earthdata) and QGIS, exposing geospatial tools to LLMs with structured metadata, confirmation hooks, and support for OpenAI, Anthropic, Google Gemini, Ollama, and more; includes the OpenGeoAgent QGIS plugin (456+ stars, MIT License)

Active4586 days ago
Python
MIT

Contrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.

Active276 days ago
Python

Deep learning-based multi-animal pose tracking and behavior classification, enabling automated quantification of social interactions and collective behavior across species (Nature Methods 2022, 2.2K+ stars)

Active6061 week ago
Python
BSD-3-Clause

Python Library for Automating Molecular Simulation: input preparation, job execution, file management, output processing and building data workflows.

Active931 week ago
Python
NOASSERTION

Foundation model for tabular data that predicts on unseen real-world tables in a single forward pass, achieving accurate small-data classification and regression without task-specific training; widely applicable to scientific datasets with limited samples (7.4K+ stars, 2022-2026)

Active7.8K1 week ago
Python
NOASSERTION

EMMO is a multidisciplinary effort to develop a standard representational framework (the ontology) for applied sciences. It is based on physics, analytical philosophy and information and communication technologies. It has been instigated by materials science to provide a framework for knowledge capture that is consistent with scientific principles and methodologies. (from GitHub)

Active901 week ago
Python
CC-BY-4.0

Technical Report 🧬

Active5.2K1 week ago
Python

Library for fast calculations of **mo**lecula**r** **fe**at**u**re**s** from 3D structures for machine learning with a focus on steric descriptors.

Active2351 week ago
Python
MIT

High-accuracy PDF→Markdown/JSON/HTML conversion, specialized for tables/formulas/code blocks with benchmark scripts

Active38.5K1 week ago
Python
Apache-2.0

OEO is a domain reference ontology for energy system modeling.

Active1571 week ago
Python
CC0-1.0

Parsers and algorithms for computational chemistry logfiles.

Active4211 week ago
Python
BSD-3-Clause

This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…

Active2131 week ago
Python

MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…

Active1.6K1 week ago
Python

linkset-automation is a set of tools to automatically generates CyTargetLinker linksets from different resources, starting with WikiPathways.

Active01 week ago
Python
Apache-2.0

For a convenient overview and download list, visit our model page for this model.

Active5781 week ago
Python

Converts Protein Data Bank structures into 3D-printable models. Each polymer chain is meshed separately and written as a named object in a single 3MF file, so a multi-material printer can assign one filament per chain. Protein chains can be rendered as a solvent-excluded surface, a cartoon, or a backbone tube; nucleic acids as a tube-and-rung form with the strands of a duplex welded at every base pair. Press-fit magnet pockets are optionally placed at chain interfaces, so a complex comes apart where its subunits actually meet. All meshes are checked for watertightness before export.

Active111 week ago
Python
MIT

PyTorch-based differentiable programming framework for physics-informed system identification, parametric constrained optimization, and model predictive control, integrating neural operators, neural ODEs, KANs, SINDy, and differentiable predictive control with 30+ tutorials (1.3k+ stars, BSD License)

Active1.4K1 week ago
Python
NOASSERTION

Robust deep learning-based segmentation of >100 anatomical structures in CT and MR images, built on nnU-Net and widely adopted in clinical radiology and surgical planning workflows (2.6K+ stars)

Active2.9K1 week ago
Python
Apache-2.0

Machine learning toolkit for many-body quantum systems, implementing neural quantum states, variational Monte Carlo, and tensor network algorithms to solve ground-state and dynamical problems in condensed matter physics and quantum chemistry (EPFL & collaborators, Nature Physics 2019/2022+, 670+ stars)

Active6911 week ago
Python
Apache-2.0

HealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.

Active221 week ago
Python

Python computational framework for analysis of single-molecule FRET data

Active11 week ago
Python
MIT

Scalable toolkit for analyzing single-cell gene expression data, including preprocessing, visualization, clustering, and trajectory inference.

Active2.5K1 week ago
Python
BSD-3-Clause

Open-source LLM-powered R&D agent framework automating data-driven AI solution building through automated research, development, and evolution; achieves top open-source performance on MLE-Bench with dual Researcher-Developer agents and supports research copilot, data mining, Kaggle, and quant R&D workflows (13.6K+ stars, MIT License, 2025-2026)

Active14.1K1 week ago
Python
MIT

Analysis of molecular dynamics trajectories.

Active7271 week ago
Python
LGPL-2.1

ProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.

Active191 week ago
Python