Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

188 of 7,050 resources

Showing 101–150

Meta FAIR's foundation model of vision, audition, and language for in-silico neuroscience, predicting fMRI brain responses to naturalistic multimodal stimuli (video, audio, text) through unified Transformer architecture mapped to the cortical surface (2026)

Active3.2K3 months ago
Jupyter Notebook
NOASSERTION

Foundation model for universal prompt-driven medical image segmentation extending SAM3 to clinical imaging, supporting 2D public benchmarks and 3D training/evaluation with text and box prompts; pretrained weights available on HuggingFace (189+ stars)

Active2203 months ago
Python
NOASSERTION

DANTE is a software tool for genotyping and characterizing tandem repeats (TRs) from both second- and third-generation sequencing data. It supports the analysis of short-read massively parallel sequencing (sr-MPS) and long-read massively parallel sequencing (lr-MPS), enabling accurate repeat characterization across a wide range of loci. A key feature of DANTE is its ability to determine genotypes at nucleotide resolution, including the characterization and phasing of complex repeat motifs. For sr-MPS data, the tool determines allele size and sequence composition of alleles for which spanning reads are generated. In addition, it identifies alleles that exceed the sequencing read length by estimating their presence from partial read evidence and supports the visualisation of the sequence composition of partial reads. For lr-MPS data, where complete repeat regions are typically sequenced, DANTE determines the allele size and sequence composition of identified alleles.

Active13 months ago
Rust
NOASSERTION

SPAdes (St. Petersburg genome assembler) is an assembly toolkit containing various assembly pipelines and the de-facto standard for prokaryotic genome assemblies.

Active9553 months ago
C++
NOASSERTION

Access to Biological Web Services from Python.

Active3423 months ago
Python
NOASSERTION

Ontology, part of the SI Reference Point, covering measurement units (SI base units and SI units with special names) and prefixes.

Active193 months ago
NOASSERTION

Sparse identification of nonlinear dynamics

Active1.9K3 months ago
Python
NOASSERTION

A swiss army knife for manipulating and editing PDB files.

Active4573 months ago
Python
NOASSERTION

Lifecycle-Aware Memory (LAM) primitive and benchmark for long-horizon research agents, achieving 65.95% mean reproduction on PaperBench and 94.66% on SurveyBench through Capital Chunk Memory (CCM) with versioned content, structural multi-hop relevance, and provenance-grounded composition; 10 peer-reviewed acceptances at FSE/ICML/TOSEM/AEI/ICoGB (1.3K+ stars)

Active1.3K4 months ago
TeX
NOASSERTION

This package provides a periodic table of the elements with support for mass, density and xray/neutron scattering information.

Active1724 months ago
Python
NOASSERTION

Minimap2 is an pairwise aligner for genomic and spliced nucleotide sequences. It can perform the assembly-to-assembly alignment, and works with gzip'd FASTQ, FASTA formats. It also finds overlaps between long-reads.

Active2.2K4 months ago
C
NOASSERTION

All-atom biomolecular structure prediction for protein-nucleic acid-small molecule-metal ion complexes, enabling accurate modeling of covalent modifications and assemblies beyond proteins (Baker Lab, Science 2024)

Active8214 months ago
Python
NOASSERTION
Active195 months ago
NOASSERTION

Benchmark quantifying end-to-end autonomous AI research abilities of LLM agents across 20 tasks from SOTA machine learning papers spanning NLP, code, math, biochemical modelling, and time series forecasting, with normalized score metrics against human SOTA and HuggingFace dataset

Active1125 months ago
Python
NOASSERTION

First physics-aligned interactive benchmark for LLM agents in engineering construction, designing rockets/cars/bridges in physics simulator with 3D spatial geometry library

Active985 months ago
Python
NOASSERTION

dominatR is an R package for quantifying and visualizing feature dominance in datasets. dominatR applies concepts drawn from physics such as center of mass and shannon's entropy to effectively visualize features (e.g. genes) that are present within a specific context or condition. The package integrates, dataframes, matrices and SummerizedExperiment objects and is able to perform common genomic normalization methods. The key aspect is the generation of plots that serve to highlight context-relevant feature dominance.

Active35 months ago
R
NOASSERTION

Arc Institute's single-cell foundation model enabling in-context learning at inference time via a novel tabular attention architecture, trained on 150M uniformly-preprocessed cells for generalizing biological effects and generating unseen cell profiles in novel contexts (2025)

Active1625 months ago
Jupyter Notebook
NOASSERTION

Benchmark evaluating AI agents on 75 curated Kaggle-style ML engineering competitions with reproducible Docker-based grading harness, human baselines, and end-to-end task lifecycle, used as a primary benchmark for autonomous ML research agents (e.g., InternAgent #1 at 36.44%)

Active1.7K5 months ago
Python
NOASSERTION

toscca is an R package to perform Thresholded Ordered Sparse Canonical Correlation Analysis (TOSCCA).

Active15 months ago
R
NOASSERTION

Dataset and benchmarking framework integrating histology and spatial transcriptomics, enabling multimodal analysis of whole-slide images with matched spatial gene expression for advancing computational pathology and tissue microenvironment research (Mahmood Lab, Harvard Medical School, 411+ stars)

Active4355 months ago
Jupyter Notebook
NOASSERTION

Unified latent diffusion transformer that jointly generates periodic crystals and non-periodic molecules, scaling to 500M parameters with SOTA results on QM9, MP20, and GEOM-DRUGS (Meta FAIR, ICML 2025, 310+ stars)

Active3185 months ago
Python
NOASSERTION

De novo assembler for single molecule sequencing reads using repeat graphs.

Idle9516 months ago
C
NOASSERTION

Baidu's open-source reproduction of AlphaFold3 in PaddlePaddle, providing pretrained weights and inference pipelines for unified biomolecular structure prediction across proteins, nucleic acids, ligands, ions, and post-translational modifications within the PaddleHelix biocomputing platform (Baidu, bioRxiv 2024)

Idle1.1K6 months ago
Python
NOASSERTION

Hybrid deep learning and alignment-based tool for identifying viruses, plasmids, and other mobile genetic elements in isolates, metagenomes, and metatranscriptomes, combining neural-network gene-content classifiers with nucleotide-sequence signatures; also performs viral taxonomic assignment, provirus detection in host genomes, and functional annotation, with precomputed databases of 200K+ viral and 1M+ plasmid genomes and web apps on Galaxy and NMDC EDGE (Berkeley Lab & DOE Joint Genome Institute, 334+ stars, actively maintained)

Idle3346 months ago
Python
NOASSERTION

Genetic variant annotation and effect prediction toolbox.

Idle3137 months ago
Java
NOASSERTION

The Generative Artificial Intelligence Delegation Taxonomy (GAIDeT) assigns identifiers to contributor roles as an extension to the Contributor Roles Taxonomy (CRediT) to support promoting transparency and accountability in academic publishing when AI contribtors are involved in research. It is operationalized in the [GAIDeT Declaration Generator](https://panbibliotekar.github.io/gaidet-declaration/), an interactive tool for researchers to disclose the delegation of tasks to generative AI (GAI) tools in accordance with the GAIDeT taxonomy.

Idle77 months ago
HTML
NOASSERTION

Multimodal deep learning framework integrating peptide-MHC protein sequence, structure, and biochemical properties to predict class-I immunogenicity for infectious disease epitopes and cancer neoepitopes with cancer-wildtype contrastive learning, enabling personalized vaccine design (Krishnaswamy Lab, Yale University)

Idle487 months ago
Python
NOASSERTION

Foundation models for genomics and transcriptomics pretrained on 3,000+ human genomes and 850+ diverse species, enabling chromatin accessibility prediction, splice site detection, and promoter classification across multiple model scales (InstaDeep, NVIDIA & TUM, Nature Methods 2023)

Idle9197 months ago
Jupyter Notebook
NOASSERTION

Euclidean neural networks for arbitrary point transformations enabling E(3)-equivariant deep learning, foundational library for building geometry-aware neural networks in molecular dynamics, materials science, and physics

Idle1.3K7 months ago
Python
NOASSERTION

Self-supervised vision foundation model for generalized structural brain MRI analysis, pretrained on ~49,000 scans from diverse datasets and generalizing across brain age prediction, dementia/MCI classification, IDH mutation detection, glioma survival prediction, time-to-stroke estimation, MR sequence classification, and brain tumor segmentation; outperforms task-specific models especially with limited training data (Mass General Brigham & Harvard Medical School, 129+ stars)

Idle1578 months ago
Python
NOASSERTION

Lightweight supervised slide foundation model with 0.9M parameters pretrained on 24K whole-slide images for pan-cancer morphological classification, achieving competitive performance with much larger self-supervised models (TITAN, GigaPath) while enabling finetuning on consumer-grade GPUs; includes standardized MIL implementations and benchmarking across 15+ classification tasks (Mahmood Lab, Harvard Medical School, 153+ stars)

Idle1528 months ago
Python
NOASSERTION

A [Jupyter](https://jupyter.org/) widget to interactively view molecular structures and trajectories.

Idle9278 months ago
Jupyter Notebook
NOASSERTION

SCENIC+ is a python package to build gene regulatory networks (GRNs) using combined or separate single-cell gene expression (scRNA-seq) and single-cell chromatin accessibility (scATAC-seq) data.

Idle2638 months ago
Jupyter Notebook
NOASSERTION

Official implementation of the second-generation fully autonomous scientific discovery system, extending the original with agentic tree search and reduced template dependency to achieve workshop-level accepted papers (6.7K+ stars, 2025)

Idle7.3K9 months ago
Python
NOASSERTION

First fully autonomous open-ended scientific discovery system with official implementation: hypothesis→experiment→writing→review simulation (13.8K+ stars, 2024)

Idle14.5K9 months ago
Jupyter Notebook
NOASSERTION

Biocaml aims to be a high-performance user-friendly library for Bioinformatics.

Idle12310 months ago
OCaml
NOASSERTION

Graph neural network operating entirely at the atomic level for protein-ligand conformational ensemble prediction and docking, generating diverse solutions through rapid stochastic denoising to model conformational heterogeneity (Baker Lab, bioRxiv 2025)

Idle26711 months ago
Python
NOASSERTION

Conversational data analysis using natural language

Idle23.8K11 months ago
Python
NOASSERTION

Structural variant and indel caller for mapped sequencing data.

Archived46712 months ago
C++
NOASSERTION

AI-assisted mutation nomination approach optimizing protein function by integrating structural and evolutionary constraints into protein inverse folding models, compatible with ProteinMPNN, LigandMPNN, ESM-IF1, and SaProt (Chinese Academy of Sciences, 359+ stars)

Idle1.1K1 year ago
Jupyter Notebook
NOASSERTION

Another list focuses on Python stuff related to Chemistry.

Idle1.4K1 year ago
NOASSERTION

HOSO is an ontology of informational entities and processes related to healthcare organizations and services.

Idle01 year ago
HTML
NOASSERTION

Cheminformatic extension for the SQLAlchemy database.

Idle401 year ago
Python
NOASSERTION

HEPRO is an ontology of informational entities and processes related to health procedures and health activities.

Idle01 year ago
HTML
NOASSERTION

SIMD C library for global, semi-global, and local pairwise sequence alignments

Idle2881 year ago
C
NOASSERTION

NIST's open-source platform for data-driven atomistic materials design, integrating DFT datasets (JARVIS-DFT), machine learning property prediction (JARVIS-ML), and a comprehensive leaderboard for benchmarking materials AI methods across the periodic table (384+ stars)

Idle4011 year ago
Python
NOASSERTION

Large-scale flow-based protein backbone generator utilizing hierarchical fold class labels for conditioning with a tailored scalable transformer architecture, enabling controllable de novo protein design (264+ stars)

Idle2761 year ago
Python
NOASSERTION

DeepSeek's open-source large language model for formal theorem proving in Lean 4, integrating informal and formal mathematical reasoning through recursive subgoal decomposition and reinforcement learning powered by DeepSeek-V3, with open weights and ProverBench evaluation (2025)

Idle1.3K1 year ago
NOASSERTION