Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain(1)
Language
License(1)
Source
Type
8 of 7,050 resources
Offline tool for cleaning tables of human gene and protein identifiers (TXT, CSV, TSV, XLSX). It maps approved symbols, aliases, previous symbols, Ensembl gene, UniProt, Entrez, RefSeq and HGNC identifiers to current HGNC approved symbols with cross-references, using a bundled HGNC snapshot. Excel date-corrupted symbols are recovered where the original is unambiguous and flagged for manual review otherwise; no input row is dropped. Each run records the tool version and HGNC release and writes a per-row audit table.
BacSelect is a reproducible resource for selecting compact, nested panels of complete bacterial genome assemblies that span genome architecture. From a defined and versioned public genome universe, BacSelect deterministically selects assemblies so users can choose a panel size from N=10–500 while retaining nestedness across scales. Versioned releases provide genome-panel manifests, downloadable assemblies, structural-coverage information, checksums and provenance for reproducible benchmarking, method development and comparative genomics.
Workflow optimized for the analysis of rare diseases, designed to detect SNVs, INDELs , CNVs and SVs in targeted sequencing data (CES/WES) and whole genome sequencing (WGS), built on Nextflow and following nf-core standards. It has an advanced variant annotation optimized for rare diseases diagnosis and discovery.
GAIn is a platform for annotating genetic variants, genomic positions, and regions with reproducible, declarative pipelines using curated Genomic Resource Repositories.
blue-crab is a tool to convert from ONT POD5 format to the community maintained SLOW5/BLOW5 format. Lossless nanopore pod5 s/blow5 file conversion.
Exact, validated excision of coordinate-defined genomic regions from transposed NEXUS matrices.
Pathogensurveillance is a population genomics pipeline for pathogen identification, variant detection, and biosurveillance. The pipeline accepts paths to raw reads for one or more organisms and creates reports in the form of an interactive HTML document. Significant features include the ability to analyze unidentified eukaryotic and prokaryotic samples, creation of reports for multiple user-defined groupings of samples, automated discovery and downloading of reference assemblies from NCBI RefSeq, and rapid initial identification based on k-mer sketches followed by a more robust multi gene phylogeny and SNP-based phylogeny.