Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

193 of 7,055 resources

Showing 151–193

NIST's open-source platform for data-driven atomistic materials design, integrating DFT datasets (JARVIS-DFT), machine learning property prediction (JARVIS-ML), and a comprehensive leaderboard for benchmarking materials AI methods across the periodic table (384+ stars)

Idle4011 year ago
Python
NOASSERTION

Large-scale flow-based protein backbone generator utilizing hierarchical fold class labels for conditioning with a tailored scalable transformer architecture, enabling controllable de novo protein design (264+ stars)

Idle2761 year ago
Python
NOASSERTION

DeepSeek's open-source large language model for formal theorem proving in Lean 4, integrating informal and formal mathematical reasoning through recursive subgoal decomposition and reinforcement learning powered by DeepSeek-V3, with open weights and ProverBench evaluation (2025)

Idle1.3K1 year ago
NOASSERTION

Unified Code for Units of Measure (UCUM) is a code system intended to include all units of measures being contemporarily used in international science, engineering, and business.

Idle1101 year ago
HTML
NOASSERTION

A project supporting the DRAO application ontology, a hierarchy of specific research domains and descriptors which imports subsets of terms from over 40 publicly-available terminologies. (from repository)

Idle21 year ago
Makefile
NOASSERTION

In silico directed evolution framework using few-shot active learning to optimize protein activities, enabling rapid protein engineering with minimal experimental data (352+ stars, 2023)

Idle3761 year ago
Python
NOASSERTION

Large language-and-vision assistant for biomedicine, instruction-tuned on GPT-4-generated biomedical multimodal instruction-following data to enable conversational visual question answering over radiology, pathology, and microscopy images, establishing open recipes for adapting general vision-language models to the biomedical domain (Microsoft Research & University of Washington, 2.2K+ stars)

Idle2.2K1 year ago
Python
NOASSERTION

GRIDSS: the Genomic Rearrangement IDentification Software Suite.

Idle2861 year ago
Java
NOASSERTION

Systematic medical RAG toolkit for question answering over PubMed, StatPearls, textbooks, and Wikipedia, supporting multiple retrievers, domain LLMs, and follow-up-query workflows for benchmarked clinical/biomedical QA (ACL Findings 2024)

Idle6001 year ago
Python
NOASSERTION

SpatialDE is a method to find spatially variable genes (SVG) from spatial transcriptomics data. This package provides wrappers to use the Python SpatialDE library in R, using reticulate and basilisk.

Idle31 year ago
R
NOASSERTION

General-purpose pathology foundation model pretrained on 100K+ diagnostic whole-slide images across 20 major tissue types, achieving state-of-the-art transfer learning across 30+ clinical tasks and serving as a universal feature extractor for digital pathology (Mahmood Lab, 722+ stars)

Idle7741 year ago
Jupyter Notebook
NOASSERTION

Vision-language pathology foundation model using contrastive learning on histopathology image-text pairs, enabling zero-shot classification, slide-level retrieval, and multimodal reasoning across diverse cancer types (Mahmood Lab, 494+ stars)

Idle5341 year ago
Python
NOASSERTION

A database system designed to store, organize, and manage large-scale nucleotide sequencing read data (like PacBio reads) for the Dazzler genome assembler

Idle361 year ago
C
NOASSERTION

A terminology for the skills necessary to make data FAIR and to keep it FAIR.

Idle171 year ago
Makefile
NOASSERTION

SKESA is a de-novo sequence read assembler for microbial genomes. It uses conservative heuristics and is designed to create breaks at repeat regions in the genome. This leads to excellent sequence quality without significantly compromising contiguity.

Idle1271 year ago
C++
NOASSERTION

Universal chart comprehension and reasoning model

Stale1362 years ago
Python
NOASSERTION

The Semantic Web for Earth and Environmental Terminology is a mature foundational ontology that contains over 6000 concepts organized in 200 ontologies represented in OWL. Top level concepts include Representation (math, space, science, time, data), Realm (Ocean, Land Surface, Terrestrial Hydroshere, Atmosphere, etc.), Phenomena (macro-scale ecological and physical), Processes (micro-scale physical, biological, chemical, and mathematical), Human Activities (Decision, Commerce, Jurisdiction, Environmental, Research).

Stale1432 years ago
Turtle
NOASSERTION

[RDKit](http://www.rdkit.org/) and [OSRA](https://cactus.nci.nih.gov/osra/) in the [Bottle](http://bottlepy.org/docs/dev/) on [Tornado](http://www.tornadoweb.org/en/stable/).

Archived502 years ago
Python
NOASSERTION

Circlator is a tool to circularize genome assemblies. It will attempt to identify each circular sequence and output a linearised version of it. It does this by assembling all reads that map to contig ends and comparing the resulting contigs with the input assembly.

Stale2592 years ago
Python
NOASSERTION

The AOPO provides classes and relationships for the semantic representation of the Adverse Outcome Pathway framework.

Stale132 years ago
Rich Text Format
NOASSERTION

k-mer counting, filtering, and graph traversal.

Stale7892 years ago
Python
NOASSERTION

NOVOPlasty - The organelle assembler and heteroplasmy caller. NOVOPlasty is a de novo assembler and heteroplasmy/variance caller for short circular genomes..

Stale2012 years ago
Perl
NOASSERTION

A VCF Parser for Python.

Stale4193 years ago
Python
NOASSERTION

Displaying sequence statistics for next-generation sequencing.

Stale253 years ago
C
NOASSERTION

lipidr an easy-to-use R package implementing a complete workflow for downstream analysis of targeted and untargeted lipidomics data. lipidomics results can be imported into lipidr as a numerical matrix or a Skyline export, allowing integration into current analysis frameworks. Data mining of lipidomics datasets is enabled through integration with Metabolomics Workbench API. lipidr allows data inspection, normalization, univariate and multivariate analysis, displaying informative visualizations. lipidr also implements a novel Lipid Set Enrichment Analysis (LSEA), harnessing molecular information such as lipid class, total chain length and unsaturation.

Stale363 years ago
R
NOASSERTION

Educational resource on performing RNA-seq analysis in the cloud using Amazon AWS cloud services. Topics include preparing the data, preprocessing, differential expression, isoform discovery, data visualization, and interpretation.

Stale1.4K3 years ago
R
NOASSERTION

A suite of algorithms for matching position weight matrices (PWM) against DNA sequences. It features advanced matrix matching algorithms implemented in C++ that can be used to scan hundreds of matrices against chromosome-sized sequences in few seconds. MOODS can also process high-order PWMs with dependencies between adjacent positions and sequence variants such as SNPs, insertions and deletions.

Stale1183 years ago
C++
NOASSERTION

Open source web framework for small molecule analysis based on Django.

Stale423 years ago
JavaScript
NOASSERTION
Stale704 years ago
Makefile
NOASSERTION

Learning nonlinear operators

Stale8424 years ago
Python
NOASSERTION

AI for chemical reaction prediction and synthesis planning

Stale4284 years ago
Python
NOASSERTION

FASTQ/A short-reads pre-processing tools: Demultiplexing, trimming, clipping, quality filtering, and masking utilities.

Stale2024 years ago
C
NOASSERTION
Stale05 years ago
NOASSERTION

This proposed vocabulary allows edges in Property Graphs (e.g Neo4j, RDF*) to be augmented with edge properties that specify ontological semantics, including (but not limited) to OWL-DL interpretations. [from GitHub]

Stale355 years ago
Makefile
NOASSERTION

Finds SNP sites from a multi-FASTA alignment file.

Stale2795 years ago
C
NOASSERTION

The Reagent Ontology (ReO) adheres to OBO Foundry principles (obofoundry.org) to model the domain of biomedical research reagents, considered broadly to include materials applied “chemically” in scientific techniques to facilitate generation of data and research materials. ReO is a modular ontology that re-uses existing ontologies to facilitate cross-domain interoperability. It consists of reagents and their properties, linking diverse biological and experimental entities to which they are related. ReO supports community use cases by providing a flexible, extensible, and deeply integrated framework that can be adapted and extended with more specific modeling to meet application needs.

Stale06 years ago
Python
NOASSERTION

Customizable pipeline for differential expression analysis with an intuitive GUI.

Stale76 years ago
Java
NOASSERTION

Flexible circular visualization of genome-associated data with BioPerl and SVG.

Stale467 years ago
Perl
NOASSERTION

Horizon chart D3-based JavaScript library for DNA data.

Stale6210 years ago
CSS
NOASSERTION