Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

260 of 7,050 resources

Showing 151–200

Automates and standardizes ligand preparation for AutoDock Vina.

Active1904 months ago
Python
Apache-2.0

Multimodal LLM-based AI agent enabling deep research in spatial transcriptomics, automating analysis and interpretation of spatial gene expression data (Harvard LiuLab, bioRxiv 2025)

Active634 months ago
Jupyter Notebook
Apache-2.0

DeepMind's neural network for ab-initio quantum chemistry, directly solving the many-electron Schrödinger equation via variational Monte Carlo with antisymmetric wavefunctions, extended to excited states (Phys. Rev. Research 2020, Science 2024)

Active8544 months ago
Python
Apache-2.0

Deployable biomedical deep-research agent blueprint combining on-prem multimodal RAG, report generation, human-in-the-loop editing, and virtual screening with MolMIM and DiffDock for drug discovery workflows (2025)

Active1444 months ago
Python
Apache-2.0

General multimodal protein design framework enabling DNA-encoding of chemistry for programmable enzyme design and diverse protein generation through diffusion-based generative modeling (190+ stars, Apache 2.0, 2026)

Active2134 months ago
Python
Apache-2.0

Open-source self-supervised vision foundation model for Earth observation by Clay Foundation (non-profit), a Masked Autoencoder ViT pretrained on multimodal satellite imagery (Sentinel-1/2, Landsat 8-9, NAIP, MODIS, LINZ DEM) with location/time embeddings, supporting classification, segmentation, change detection, similarity search, and few-shot downstream geospatial tasks (Apache 2.0, v1.5 2024-2025)

Active6134 months ago
Python
Apache-2.0

Fast, all-atom SE(3)-equivariant diffusion model for protein design achieving state-of-the-art performance on unconditional generation, motif scaffolding, and binder design while retaining the computational efficiency of equivariant architectures (bioRxiv 2026)

Active1355 months ago
Python
Apache-2.0

2D interactive visualization in Jupyter.

Active3.7K5 months ago
TypeScript
Apache-2.0

General-purpose RNA language model with 650M parameters pretrained on 36M non-coding RNA sequences, achieving strong generalization on structure prediction tasks including secondary structure prediction, splice-site prediction, mean ribosome loading, and ncRNA classification (lbcb-sci, 165+ stars, Apache-2.0)

Active1735 months ago
Python
Apache-2.0

Incremental knowledge graph construction using LLMs with entity extraction and Neo4j visualization

Active9575 months ago
Python
Apache-2.0

Protein structure prediction

Active14.9K5 months ago
Python
Apache-2.0

FutureHouse's end-to-end scientific discovery multi-agent system orchestrating literature search (Crow/Falcon) and data analysis (Finch) agents, first AI-generated drug discovery identifying ripasudil as novel dry AMD therapeutic (2025)

Active6815 months ago
Python
Apache-2.0

Descriptor library containing a variety of fingerprinting techniques, including the Smooth Overlap of Atomic Positions (SOAP).

Active4755 months ago
C++
Apache-2.0

Fully autonomous medical image segmentation research system that generates complete manuscripts end-to-end from datasets with zero human intervention, beating strongest baselines on 24 of 31 datasets and achieving T1-T2 tier manuscript quality in double-blind evaluations (USTC & Shanghai AI Lab, 2026)

Active3675 months ago
Python
Apache-2.0

First large vision-language assistant for gigapixel whole-slide pathology image understanding, released with the SlideInstruction dataset and SlideBench benchmark (uni-medical, Apache 2.0, 2025)

Idle1266 months ago
Python
Apache-2.0

Allen Institute for AI's global geospatial foundation model for satellite imagery analysis, enabling large-scale mapping of buildings, wind turbines, trees, and land cover from Sentinel-2 data with open-source weights and inference tools (2024)

Idle2856 months ago
Python
Apache-2.0

Toolkit for linearizing academic PDFs into LLM-ready text with high accuracy and structure preservation, optimized for scientific literature extraction

Idle19.5K6 months ago
Python
Apache-2.0

End-to-end semi-automated scientific discovery system that designs, iterates, and analyzes code-based experiments via LLM-as-a-mutator over scientific articles and code examples; auto-creates, runs, and debugs experiment code in containers and writes meta-analysis reports (339+ stars, Apache 2.0)

Idle3486 months ago
Python
Apache-2.0

Automated code generation from machine learning research papers into runnable implementations (4.5K+ stars, 2025)

Idle4.9K6 months ago
Python
Apache-2.0

Target-aware peptide design framework that treats receptor sequence and structure as context via multimodal adapter tuning of protein language models (ESMC + ProteinMPNN features), with reinforcement-learning-based 3D dynamic feedback (ESMFold structure evaluation) to suppress unrealistic peptide conformations; ships with a systematic assessment pipeline covering peptide-target affinity, structure quality, physicochemical properties, diversity, and novelty (142+ stars, Apache 2.0, 2026)

Idle1426 months ago
Jupyter Notebook
Apache-2.0

Bi-directional DNA language model based on the Mamba state space architecture, enabling efficient long-range genomic sequence modeling with linear-time complexity and built-in reverse-complement equivariance; achieves strong performance on chromatin accessibility, enhancer, and promoter prediction benchmarks (Stanford & UC Berkeley, 500+ stars)

Idle2526 months ago
Python
Apache-2.0

This tutorial aims to illustrate the process of analyzing a membrane molecular dynamics (MD) simulation, step by step, using the BioExcel Building Blocks (biobb)

Idle17 months ago
Jupyter Notebook
Apache-2.0

This tutorial aims to illustrate the process of checking a molecular structure before using it as an input for a Molecular Dynamics simulation, step by step, using the BioExcel Building Blocks (biobb).

Idle07 months ago
HTML
Apache-2.0

This BioExcel Building Blocks library (BioBB) workflow provides a pipeline to setup DNA structures for the Ascona B-DNA Consortium (ABC) members. It follows the work started with the NAFlex tool to offer a single, reproducible pipeline for structure preparation, ensuring reproducibility and coherence between all the members of the consortium.

Idle17 months ago
HTML
Apache-2.0

This tutorial aims to illustrate the process of generating protein conformational ensembles from 3D structures using Coarse-Grained tools from the FlexServ server and analysing its molecular flexibility

Idle37 months ago
HTML
Apache-2.0

This tutorial aims to illustrate the process of computing a conformational transition between two known structural conformations of a protein, step by step, using the BioExcel Building Blocks (biobb).

Idle07 months ago
HTML
Apache-2.0

This tutorial aims to illustrate the process of protein-ligand docking, step by step, using the BioExcel Building Blocks library (biobb).

Idle77 months ago
HTML
Apache-2.0

This tutorial aims to illustrate the process of computing classical molecular interaction potentials from protein structures step by step, using the BioExcel Building Blocks library (biobb)

Idle17 months ago
HTML
Apache-2.0

This tutorial aims to illustrate how to compute a fast-growth mutation free energy calculation, step by step, using the BioExcel Building Blocks library (biobb). The particular example used is the Staphylococcal nuclease protein (PDB code 1STN), a small, minimal protein, appropriate for a short tutorial.

Idle57 months ago
HTML
Apache-2.0

The initial focus of the GS1 Web Vocabulary is consumer-facing properties for clothing, shoes, food beverage/tobacco and properties common to all products. [from homepage]

Idle528 months ago
Apache-2.0

Foundation model for joint segmentation, detection, and recognition of biomedical objects across nine imaging modalities, with v2 introducing BoltzFormer architecture for end-to-end 3D inference (Microsoft, Nature Methods 2025)

Idle7008 months ago
Python
Apache-2.0

DNA sequence analysis

Idle7788 months ago
Python
Apache-2.0

DeepMind's Olympiad-level geometry theorem prover combining neural language model with symbolic deduction engine, AlphaGeometry2 solves 84% of IMO geometry problems (42/50) at gold-medalist level (Nature 2024)

Idle4.9K8 months ago
Python
Apache-2.0

Standard data-centric AI package for data quality and machine learning, automatically detecting label errors, outliers, and dataset issues to improve scientific dataset reliability and model performance (11K+ stars, MIT License)

Idle11.7K8 months ago
Python
Apache-2.0

Efficient foundation model and benchmark for multi-species genome understanding with context-aware nucleotide representations, improving upon DNABERT for diverse genomic task transfer learning (UIUC MAGICS Lab, 484+ stars)

Idle5129 months ago
Shell
Apache-2.0

Fast, modular, and accurate de novo design of protein binders based on the Protenix foundation model, achieving 17-82% nanomolar hit rates across diverse targets with 2-6× improvement over prior methods like AlphaProteo and RFdiffusion (229+ stars, Apache 2.0)

Idle2569 months ago
Python
Apache-2.0

Trainable, memory-efficient PyTorch reproduction and retraining of AlphaFold2 providing new insights into its learning dynamics and out-of-distribution generalization; widely used as the open-source AlphaFold2 backbone underpinning many downstream protein structure prediction and design pipelines (Columbia AlQuraishi Lab & OpenFold Consortium, Nature Methods 2024)

Idle3.4K9 months ago
Python
Apache-2.0

Full spaCy pipeline and models for scientific/biomedical documents, enabling named entity recognition, abbreviation resolution, and UMLS linking for scientific literature mining (1.9K+ stars, Apache 2.0)

Idle2K10 months ago
Python
Apache-2.0

Autonomous multi-agent research loop for model architecture discovery that ran 1,773 experiments over 20,000 GPU hours and produced 106 state-of-the-art linear-attention architectures, surpassing human-designed baselines including Mamba2 and DeltaNet (1.1K+ stars, Apache 2.0)

Idle1.2K10 months ago
Python
Apache-2.0

100M-parameter foundation model pretrained on 50M+ human single-cell transcriptomes covering ~20,000 genes, achieving SOTA on gene expression enhancement, drug response and perturbation prediction (Nature Methods 2024)

Idle43110 months ago
Jupyter Notebook
Apache-2.0

Teaching Large Language Models the Language of Biology through single-cell transcriptomics (ICML 2024)

Idle87811 months ago
Jupyter Notebook
Apache-2.0

First versatile medical reasoning agent for chest X-ray interpretation, dynamically integrating state-of-the-art CXR analysis tools and multimodal LLMs into a unified framework; introduces ChestAgentBench with 2,500 complex medical queries across 7 categories (bowang-lab, 1.1K+ stars)

Idle1.2K11 months ago
Python
Apache-2.0

A library for computational chemistry (DFT) for input file generation, data extraction, method screening and analysis.

Idle2211 months ago
Python
Apache-2.0

Generalist foundation model and database for open-world medical image segmentation, enabling universal segmentation of diverse anatomical structures and pathologies with zero-shot generalization to unseen tasks and modalities (Nature Biomedical Engineering 2025)

Idle911 year ago
Python
Apache-2.0

Automated and rigorous experiments using AI agents for scientific discovery

Idle3681 year ago
Python
Apache-2.0

Open-ended self-improving agent that iteratively rewrites its own codebase and empirically validates each mutation on coding benchmarks (SWE-bench, Polyglot), demonstrating open-ended evolution where agents improve their ability to improve themselves, diverging into a population of diverse specialists (arXiv 2505.22954, 2.3K+ stars, Apache 2.0, 2025)

Idle2.4K1 year ago
Python
Apache-2.0

Retrieval-augmented LM synthesizing scientific literature from 45M papers with human-expert-level citation accuracy, outperforming GPT-4o by 5% on ScholarQABench (Nature 2026, UW & Ai2)

Idle1.6K1 year ago
Python
Apache-2.0

Deep Graph Library for scalable deep learning on graphs, powering molecular modeling, materials discovery, protein interaction networks, and scientific knowledge graph learning across PyTorch, TensorFlow, and MXNet backends (14K+ stars)

Idle14.3K1 year ago
Python
Apache-2.0

Family of diffusion protein language models demonstrating versatile generative and predictive capabilities for protein sequences and structures, including multimodal co-generation, conditional folding, inverse folding, motif scaffolding, and representation learning, with open pretrained weights and training scripts (327+ stars, ICML 2024, ICLR 2025, ICML 2025 Spotlight)

Idle3441 year ago
Python
Apache-2.0

Public release of Profluent's ProGen3 protein language model family, including PMC-15B supporting sequence- and structure-conditioned generation for protein design, zero-shot fitness prediction, and antibody engineering with state-of-the-art performance on fitness and docking benchmarks (114+ stars, Apache 2.0)

Idle1141 year ago
Python
Apache-2.0