Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type
1,191 of 7,068 resources
Showing 201–250
BacSelect is a reproducible resource for selecting compact, nested panels of complete bacterial genome assemblies that span genome architecture. From a defined and versioned public genome universe, BacSelect deterministically selects assemblies so users can choose a panel size from N=10–500 while retaining nestedness across scales. Versioned releases provide genome-panel manifests, downloadable assemblies, structural-coverage information, checksums and provenance for reproducible benchmarking, method development and comparative genomics.
Differentiable tokamak core transport simulator for fusion energy research, coupling PDE solvers with JAX auto-differentiation and neural-network surrogates for fast forward modelling, pulse-design, and trajectory optimization (Google DeepMind, Apache 2.0)
Developer toolkit for accelerating training and inference for AI in chemistry and material science, providing optimized GPU-accelerated workflows for molecular and materials machine learning (NVIDIA, 2026)
An ontology written in OWL 2 DL to enable characterization of the five attributes of an online journal article - peer review, open access, enriched content, available datasets and machine-readable metadata.
Universal machine learning interatomic potential for atomistic simulation of materials, molecules, and biomolecules across the periodic table, with open-source pretrained models and inference tools (Orbital Materials, 2024-2025)
ECMWF's open-source machine-learning Earth system model developed by the WeatherGenerator Consortium with NVIDIA, trained on reanalyses, forecast data, and diverse observations across atmosphere, ocean, and land to provide a robust multi-scale model of Earth system dynamics; the first released version (v0.1, trained on ERA5) demonstrates global probabilistic forecasting skill on par with established AI models, with open training framework and config-driven multi-dataset ingestion pipeline (Apache 2.0)
An ontology based on PRO for describing the contributions that may be made, and the roles that may be held by a person with respect to a journal article or other publication (e.g. the role of article guarantor or illustrator).
An ontology for describing the steps in the workflow associated with the publication of a document or other publication entity.
Generalist autonomous research agent that grows a hypothesis tree to optimize any measurable task, beating Claude Code and Codex by 2.5× on the same compute budget across BrowseComp, Terminal-Bench 2.0, math reasoning, and MLE-Bench Lite; supports native CLI, keyless Claude Code/Codex integration, and an MCP tool server (RUC-NLPIR, 866+ stars, Apache 2.0, 2026)
High-performance symbolic regression for discovering interpretable scientific equations from data, multi-population evolutionary search with Python/Julia backend, widely used in physics and astronomy (Cambridge, NeurIPS 2023)
Language agent gymnasium for challenging scientific tasks including DNA manipulation, literature search, and protein engineering
Deep learning package for many-body potential energy representation and molecular dynamics, achieving quantum-mechanical accuracy with classical MD efficiency (DeepModeling, Gordon Bell Prize 2020, 1.9k+ stars)
Democratizing AI scientists by transforming any LLM into research systems with 600+ scientific tools (Harvard MIMS)
A classification of subjects in Hochschule (universities of applied sciences)
Fits second-order autoregressive AR(2) models to gene expression time series and reports the eigenvalue modulus |lambda|, a single statistic quantifying temporal persistence: how strongly a gene's recent past constrains its next value. Ranks genes into a clock/target/background hierarchy and reports correlation length, half-life and root type (real or complex) per gene.
Robust, lightweight infrastructure for multi-agent autonomous self-evolution, built for autoresearch; agents run in isolated git worktrees, share knowledge through a common state directory, and are scored by a grader daemon; natively integrated with Claude Code, Codex, Cursor Agent, OpenCode, and Kiro (672+ stars, Apache 2.0)
Predicts transcription factor binding sites in up to 316 vertebrate species by scoring JASPAR matrices against Ensembl promoter sequences and combining the match with seven contextual experimental datapoints, including evolutionary conservation, CAGE-defined transcription start sites, eQTLs, ChIP-seq peaks, ATAC-seq accessibility, DNase footprints and gene expression correlation, into a single score per site.
The EMI ontology is used to structure spectrum annotation provenance by reusing the PROV-O ontology (a W3C recommendation) and sample and observation data by applying the SOSA ontology. EMI reuses the SOSA ontology as a data schema for struturing the Sample and Observation data. SOSA (Sensor, Observation, Sample, and Actuator) is a subset of SSN (Semantic Sensor Network Ontology) that is a W3C recommendation. [from homepage]
Benchmark evaluating AI agents on complex real-world scientific workflows in terminal environments across life, physical, earth, and mathematical sciences; featured on model cards for Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro (200+ stars, Apache 2.0)
Text-space optimizer that treats agent skill documents as trainable parameters for frozen LLMs, using scored rollouts and held-out validation gates to iteratively improve reusable natural-language skills; includes SkillOpt-Sleep for nightly self-evolution and improves accuracy across Claude Code, Codex, Copilot, and direct-chat harnesses, making it a meta-tool for evolving scientific agent skill workflows (15.5K+ stars, MIT License, PyPI)
A Simulation Tool for Fractured and Deformable Porous Media.
TransformerEngine-accelerated checkpoints and training recipes for scaling biological foundation models (ESM-2, AMPLIFY, Geneformer, CodonFM) from single-GPU prototyping to multi-node FSDP training with FP8/MXFP8/NVFP4 precision, compatible with PyTorch, HF Accelerate, and PyTorch Lightning, plus sparse-autoencoder interpretability tools for biological foundation models (851+ stars, 2026)
Foundation AutoResearch Operating System: blueprint-driven runtime for orchestrating AI research workflows from idea generation and experiments to paper writing and peer review (OpenNSWM-Lab, 2.4K+ stars, 2026)
Flow-based generative model for atomistic protein binder design with test-time optimization, SOTA on binder benchmarks (ICLR 2026 Oral, NVIDIA)
MITE (Minimum Information about a Tailoring Enzyme) is a data repository and associated data standard designed to capture the reaction- and substrate-specificities of tailoring enzymes. Community-driven and fully expert-reviewed, it represents enzymatic reactions using reaction SMARTS and links to established resources such as UniProt, NCBI GenPept, Rhea, and MIBiG. MITE serves as a knowledgebase for enzyme and pathway annotation, in silico biosynthesis, and machine learning applications.
NVIDIA and King's College London's open-source AI toolkit for healthcare imaging, providing foundational frameworks for medical image annotation (MONAI Label), training (MONAI Core), and deployment (MONAI Deploy) across radiology, pathology, and endoscopy (8K+ stars, Apache 2.0)
ReviewAid is an open-source AI-assisted tool for full-text screening and data extraction in systematic reviews. It supports evidence synthesis workflows by using large language models to classify articles according to user-defined PICO criteria and extract structured information from full-text publications. ReviewAid is designed as a supplementary reviewer rather than a replacement for human judgement. It aims to reduce manual workload, improve consistency, and assist researchers during screening and data extraction while maintaining human oversight throughout the evidence synthesis process.
Family of operational-quality open weather models from DeepMind and Google Research, including WeatherNext Graph (deterministic GNN medium-range forecasting, published as GraphCast), WeatherNext Gen (diffusion ensemble, published as GenCast), WeatherNext 2 (state-of-the-art global medium-range and cyclone forecasting skillful beyond 15 days, operational at 0.25° resolution), and WeatherNext Cyclones (breakthrough tropical cyclone track forecasting, Nature 2026); official code and open weights (Apache 2.0, 7.6K+ stars)
High-throughput PubChem client for batch queries with caching, validation, rate-limit-aware retries, and a simple CLI.
Freely available tools for biological computing in Python, with included cookbook, packaging and thorough documentation. Part of the [Open Bioinformatics Foundation](http://open-bio.org/). Contains the very useful [Entrez](https://biopython.org/DIST/docs/api/Bio.Entrez-module.html) package for API access to the NCBI databases.
A package to 'build' collections of materials properties from the output of computational materials calculations.
METPO (Microbial Ecophysiological Trait and Phenotype Ontology) provides standardized terms for describing microbial phenotypes, growth characteristics, and culture conditions. It includes classes for growth media, temperature tolerances, pH tolerances, and relationships like "grows in" and "does not grow in".
First fully customizable open-source multiagent framework automating complete research lifecycle from idea conception to LaTeX papers with dynamic workflows
An ontology that permits the number of in-text citations of a cited source to be recorded, together with their textual citation contexts, along with the number of citations a cited entity has received globally on a particular date.
An ontology meant to define bibliographic records, bibliographic references, and their compilation into bibliographic collections and bibliographic lists, respectively.
An ontology that provides a structured vocabulary written of document components, both structural (e.g., block, inline, paragraph, section, chapter) and rhetorical (e.g., introduction, discussion, acknowledgements, reference list, figure, appendix).
An ontology that enables characterization of the nature or type of citations, both factually and rhetorically.
JCVI is a versatile toolkit for comparative genomics analysis. It is a collection of Python libraries to parse bioinformatics files, or perform computation related to assembly, annotation, and comparative genomics.
Official Jupyter extension with `%%ai` magic commands and sidebar chat assistant, connecting multiple model providers and local inference
First large vision-language assistant for gigapixel whole-slide pathology image understanding, released with the SlideInstruction dataset and SlideBench benchmark (uni-medical, Apache 2.0, 2025)
Library for fast calculations of **mo**lecula**r** **fe**at**u**re**s** from 3D structures for machine learning with a focus on steric descriptors.
Comprehensive collection of 125+ ready-to-use scientific skill modules for Claude AI across bioinformatics, cheminformatics, clinical research, ML, and materials science
Converts Protein Data Bank structures into 3D-printable models. Each polymer chain is meshed separately and written as a named object in a single 3MF file, so a multi-material printer can assign one filament per chain. Protein chains can be rendered as a solvent-excluded surface, a cartoon, or a backbone tube; nucleic acids as a tube-and-rung form with the strands of a duplex welded at every base pair. Press-fit magnet pockets are optionally placed at chain interfaces, so a complex comes apart where its subunits actually meet. All meshes are checked for watertightness before export.
Open-source hybrid scientific research agent and workbench replicating Claude Science, combining JSON tool orchestration with persistent Python/R Code-as-Action kernels, 604 bundled science skills, MCP connectors, sandboxed local execution, and multi-provider LLM support for end-to-end scientific workflows (PKU–YuanKong Intelligence, 377+ stars, MIT License, 2026)
Comprehensive Claude Code skill suite covering the full academic pipeline from deep research and paper writing to multi-perspective peer review, revision, and finalization; features multi-agent teams, PRISMA systematic review, style calibration, claim-level citation audits, integrity gates, and human-in-the-loop safeguards (38K+ stars, CC BY-NC 4.0, 2026)