Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

1,191 of 7,068 resources

Showing 101–150

Non-invasive decoding of typed sentences from MEG and EEG brain recordings using a convolutional encoder, transformer, and character-level language model; official code for the Nature Neuroscience paper and Meta blog post on brain-AI communication (Meta FAIR, 894+ stars, CC BY-NC 4.0, 2026)

Active9372 weeks ago
Python
NOASSERTION

Create MSP files containing the isotopic patterns for given molecules with given adducts. The tool is based on enviPat and the RforMassSpectrometry toolbox.

Active152 weeks ago
Python
MIT

🩺 HuatuoGPT-3-27B 🏠 GitHub | 📄 Paper

Active1202 weeks ago
Python

🩺 HuatuoGPT-3-9B 🏠 GitHub | 📄 Paper

Active2862 weeks ago
Python

Scalable toolkit for analyzing single-cell gene expression data, including preprocessing, visualization, clustering, and trajectory inference.

Active2.6K2 weeks ago
Python
BSD-3-Clause

HuatuoGPT-3-Grader-8B GitHub | Paper

Active6552 weeks ago
Python

First agentic LLM for autonomous data science with end-to-end pipeline from data to analyst-grade reports

Active4.7K2 weeks ago
Python
MIT

A toolkit for visualizations in materials informatics.

Active3352 weeks ago
Python
MIT

Open-source LLM-powered R&D agent framework automating data-driven AI solution building through automated research, development, and evolution; achieves top open-source performance on MLE-Bench with dual Researcher-Developer agents and supports research copilot, data mining, Kaggle, and quant R&D workflows (13.6K+ stars, MIT License, 2025-2026)

Active14.7K2 weeks ago
Python
MIT

Unified Python framework for bulk, single-cell, and spatial RNA-seq multi-omics analysis with deep learning deconvolution (VAE) and graph neural networks, bridging Bindea, Bindea, scanpy and squidpy ecosystems (Nature Communications 2024)

Active1.2K2 weeks ago
Python
GPL-3.0

Scalable genomic analysis.

Active1.1K2 weeks ago
Python
MIT

Unified Python framework for extracellular electrophysiology, standardizing interfaces to 10+ ML-based spike sorting algorithms including Kilosort for reproducible neural spike sorting workflows (792+ stars, actively maintained)

Active8492 weeks ago
Python
MIT

StaphScope is an automated, locally-executable computational pipeline designed specifically for comprehensive Staphylococcus aureus genomic surveillance. It addresses the critical bottleneck in MRSA research by integrating seven essential genotyping methods into a single, cohesive workflow.

Active362 weeks ago
Python
MIT

Developed by

Active1882 weeks ago
Python

English | 简体中文

Active6732 weeks ago
Python

Ultra-fast extraction of predefined clinical variables from free-text clinical notes.

Active982 weeks ago
Python

Python package for segmenting geospatial data with the Segment Anything Model (SAM), enabling zero-shot object segmentation in satellite and aerial imagery for remote sensing and Earth observation (MIT, 4k+ stars)

Active4.1K2 weeks ago
Python
MIT

Research coding benchmark curated by scientists with 338 subproblems across 16 subdomains (physics, math, materials, biology, chemistry), evaluating LLMs on realistic scientific programming tasks with gold-standard solutions (NeurIPS 2024)

Active2312 weeks ago
Python
Apache-2.0

PyTorch domain library for geospatial deep learning providing standardized datasets, samplers, transforms, and pre-trained models for remote sensing, land cover mapping, and environmental monitoring (Microsoft, 4K+ stars)

Active4.2K2 weeks ago
Python
MIT

dadi is a bioinformatics tool for inferring demographic history and selection from genetic data using diffusion approximations, offering speed and flexibility in modeling population dynamics. It supports up to three populations with customizable parameters and provides efficient computational performance.

Active82 weeks ago
Python
NOASSERTION

Python library for blazing-fast genomic interval operations and genomic file formats I/O on Polars DataFrames

Active2002 weeks ago
Python
Apache-2.0

PseudoScope is an automated, locally-executable computational pipeline designed specifically for comprehensive Pseudomonas aeruginosa genomic surveillance. It integrates seven essential analysis modules into a single, cohesive workflow: FASTA QC (assembly quality metrics), MLST (Oxford scheme), PAST serotyping (O-antigen typing), AMRFinderPlus (antimicrobial resistance gene detection), ABRicate (multi-database screening for resistance, virulence, plasmids, biocides), Ultimate Reporter (gene-centric integration with interactive HTML), and Visualisation Dashboard (publication-ready interactive plots including PCA, networks, boxplots). PseudoScope runs entirely locally (or on HPC clusters), protects data privacy, and produces beautiful interactive reports in minutes.

Active102 weeks ago
Python
MIT

Kleboscope is an automated, locally‑executable computational pipeline designed specifically for comprehensive Klebsiella pneumoniae genomic surveillance. It addresses the growing threat of multidrug‑resistant and hypervirulent K. pneumoniae by integrating eight essential analysis modules into a single, cohesive workflow. Kleboscope offers two complementary report views: Gene‑centric – each gene is shown with all genomes that contain it, together with its frequency, enabling rapid cross‑genome pattern discovery; and Sample‑centric – each isolate gets its own interactive box with typing badges (MLST, K‑locus, O‑locus, hypervirulence), per‑database tables (AMR, Virulence, BACMET, Plasmids), and full mutation details – perfect for clinical reports and patient‑level investigations.

Active62 weeks ago
Python
MIT

Python library to train, interpret, and apply deep learning models to DNA sequences, providing a unified framework for regulatory genomics with support for CNN and transformer architectures, variant effect prediction, and attribution analysis (325+ stars)

Active3652 weeks ago
Python
MIT

AcinetoScope is an automated, comprehensive bioinformatics pipeline designed specifically for the genomic analysis of Acinetobacter baumannii, a WHO Critical Priority pathogen responsible for devastating hospital-acquired infections. It integrates seven analysis types (MLST, ABRicate, AMRFinder, Kaptive 3, APT, PlasmidFinder, and mutation detection) into a single automated workflow — from FASTA to actionable insights. The pipeline offers both gene-centric and sample-centric reporting, dynamic grouping by typing, and is optimised for HPC, cloud, and container environments.

Active203 weeks ago
Python
MIT

Ensemble of automated QM workflows that can be run through jupyter notebooks, command lines and yaml files.

Active1333 weeks ago
Python
MIT

Neuro-imaging file formats.

Active7933 weeks ago
Python
NOASSERTION

Family of codon-resolution language models trained on 130 million protein-coding sequences from over 20,000 species, enabling cross-species gene expression prediction and codon-level functional genomics (2025)

Active933 weeks ago
Python
Apache-2.0

Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs (460+ stars, 2024)

Active4973 weeks ago
Python
MIT

Controllable foundation model for general and specialized biomolecular structure prediction across proteins, nucleic acids, and complexes, featuring a public web server for interactive prediction workflows (IntelliGen AI, 223+ stars, Apache 2.0, 2025)

Active2303 weeks ago
Python
Apache-2.0

Meta's comprehensive ML ecosystem for materials/chemistry with 118M+ DFT calculations, EquiformerV2 models achieving top Matbench Discovery performance

Active2.3K3 weeks ago
Python
NOASSERTION

Convert AMBER forcefields from ANTECHAMBER to GROMACS format.

Active2633 weeks ago
Python
GPL-3.0

Multi-agent system automatically transforming research papers into interactive AI agents with MCP server generation, tutorial auto-detection, and benchmark extraction (2.2K+ stars, MIT License, 2025)

Active3.5K3 weeks ago
Python
MIT

Microsoft's generative model for sampling protein equilibrium conformations 100,000× faster than MD simulations, predicting domain motions, local unfolding and cryptic binding pockets on a single GPU (Science 2025)

Active8783 weeks ago
Python
MIT

A library for processing, analyzing and modeling spectroscopic data.

Active1843 weeks ago
Python
NOASSERTION

The HGVS Nomenclature is an internationally-recognized standard for the description of DNA, RNA and protein sequence variants. It is used to convey variants in clinical reports and to share variants in publications and databases. The HGVS Nomenclature is administered by the [HGVS Variant Nomenclature Committee (HVNC)](https://hgvs-nomenclature.org/stable/hvnc/) under the auspices of the [Human Genome Organization (HUGO)](https://hugo-int.org/).

Active133 weeks ago
Python
MIT

Curated library of 550+ medical research agent skills spanning evidence insights, protocol design, omics/clinical data analysis, and academic writing; each skill is reviewed through MedSkillAudit and compatible with Claude Code, Codex, Open Code, OpenClaw, and SKILL.md-compatible agents (AIPOCH, 1.2K+ stars, MIT License, 2026)

Active1.9K3 weeks ago
Python
MIT

IBM's open foundation model family for materials and chemistry, covering SMILES, SELFIES, molecular graphs, 3D atom positions, and electron density grids, with a unified toolkit for representation learning and downstream prediction/generation (Apache 2.0, 2024-2025)

Active3223 weeks ago
Python
Apache-2.0

Some IDs may represent experiment sets, e.g. https://www.mavedb.org/#/experiment-sets/urn:mavedb:00000011 Others represent genomic regions (specifically deep mutational scans thereof) e.g. https://www.mavedb.org/#/experiment-sets/urn:mavedb:00000011-a

Active183 weeks ago
Python
AGPL-3.0

Agricultural machine learning platform

Active3463 weeks ago
Python
Apache-2.0

Deep learning-based multi-animal pose tracking and behavior classification, enabling automated quantification of social interactions and collective behavior across species (Nature Methods 2022, 2.2K+ stars)

Active6143 weeks ago
Python
BSD-3-Clause

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active1.1K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

The ESMC scaling-study checkpoints are being released to support reproducibility of the findings in our paper, please refer to the paper and github for details. Please use the ESMC model for research work. ESMC is a state-of-the-art protein language model trained on billions of protein sequences…

Active19.8K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active528.3K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active8.4K3 weeks ago
Python

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active22K3 weeks ago
Python