Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source(1)
Type
23 of 6,590 resources
PathForge is a modular benchmarking framework for multiple instance learning in computational pathology. It supports whole slide image feature extraction, HDF5 artifact generation, tile overviews, benchmarking, pipeline optimization, classification, regression, survival and retrieval tasks, and support for model inference and visualization.
Composable computational-science methodology skills for AI research agents emphasizing pre-registration, reproducible workspaces, and red-team review to guard against p-hacking and HARKing; zero third-party dependencies and runs with any agent harness plus a POSIX shell (281+ stars, MIT License, 2026)
PathBench-MIL is a comprehensive, flexible benchmarking/AutoML framework for multiple instance learning in histopathology. PathBench-MIL is expected to be deprecated and replaced by PathForge.
This tool estimates the completeness of KEGG pathway modules from the presence or absence of KEGG orthologues (KOs)
Terminal AI coding assistant with a built-in math formalization engine that converts plain-language math problems into Lean 4 theorems and attempts formal proofs; bundles a local Lean toolchain and WebUI for interactive mathematical reasoning (math-ai-org, 582+ stars, 2026)
Databank of optimised macromolecular structures. PDB-REDO entries are refined, rebuilt and validated with one consistent protocol using the equivalent entry in the Protein Data Bank and its experimental data. PDB-REDO entries typically have higher structural quality and a better fit to the experimental data.
blue-crab is a tool to convert from ONT POD5 format to the community maintained SLOW5/BLOW5 format. Lossless nanopore pod5 s/blow5 file conversion.
CebraEM is a bioinformatics tool for analyzing and processing large-scale imaging data, providing a pipeline for segmentation, annotation, and analysis with support for both Linux and Windows environments. It includes modules for core functionality, annotation, and network analysis, requiring specific dependencies and a conda environment for execution.
Efficient foundation model and benchmark for multi-species genome understanding with context-aware nucleotide representations, improving upon DNABERT for diverse genomic task transfer learning (UIUC MAGICS Lab, 484+ stars)
The Data Science Ontology is a research project of IBM Research AI and Stanford University Statistics. Its long-term objective is to improve the efficiency and transparency of collaborative, data-driven science.
The midlevel energy ontology (MENO) is a BFO-based midlevel ontology. It comprises the concepts for energy qualities, energy-based dispositions and energy-driven transformation and transfer processes and their interrelations. It has the goal to provide an upper level structure for these concepts for energy-related domain ontologies.
Syntax Highlighting for Computational Biology file formats (SAM, VCF, GTF, FASTA, PDB, etc...) in vim/less/gedit/sublime.
Subvolume processing scripts with the TOM toolbox is a collection of scripts form a pipeline for subvolume alignment and averaging of electron cryo-tomography data.
A cross-system scripting language for working with big data pipelines in computer systems of different sizes and capabilities.
Virtual machine with all software and sample data to run 3D-e-Chem Knime workflows