Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type(1)
37 of 6,565 resources
A collection of object-oriented software tools for problems involving chemical kinetics, thermodynamics, and transport processes.
Universal molecular toolkit that can be used for molecular fingerprinting, substructure search, and molecular visualization written in C++ package, with Java, C#, and Python wrappers.
An ultrafast protein aligner for `blastp` and `blastx` like searches.
High-performance molecular simulation toolkit
A software package for estimating gene and isoform expression levels from RNA-Seq data.
LLMs as copilots for theorem proving in Lean 4, exposing native tactics (`suggest_tactics`, `search_proof`, `select_premises`) that embed language model inference and premise retrieval directly inside the Lean proof environment, supporting local CTranslate2/CUDA inference as well as remote model APIs for interactive and automated proof search (Caltech & NVIDIA, NeurIPS 2024, 1.2K+ stars)
Structural variant discovery by integrated paired-end and split-read analysis.
Oxford Nanopore's official deep-learning basecaller for nanopore sequencing, converting raw electrical signals into DNA/RNA sequences with integrated modified-base (methylation) detection and efficient CPU/GPU inference; foundational tool for long-read genomics, epigenetics, and real-time sequencing analysis (nanoporetech, 846+ stars, actively maintained)
The modern C++ library for sequence analysis.
Genome mapping and spliced alignment of cDNA or amino acid sequences
Deep learning framework for molecular docking extending AutoDock Vina with convolutional neural network scoring functions, achieving superior virtual screening enrichment and pose prediction across diverse target classes; widely adopted in pharmaceutical structure-based drug design (J. Cheminformatics, 915+ stars, actively maintained)
A single molecule sequence assembler for genomes large and small.
SPAdes (St. Petersburg genome assembler) is an assembly toolkit containing various assembly pipelines and the de-facto standard for prokaryotic genome assemblies.
An ultrafast and memory-efficient tool for aligning sequencing reads to long reference sequences.
A haplotype-resolved assembler for accurate Hifi reads.
A small <720Kb C++ windows utility. That allows you to load Ancestry, 23andMe, FTDNA, or Genes for Good RAW DNA files search them, merge them. covert them to Ancestry format. But also create files from peer reviewed publications to compare with you loaded data to give your genetic disposition for the condition you have entered the data for an statistical risk if OR values are included. Included with the program are example files for Type 2 Diabetes risk factors. (As I have type 2 Diabetes so I could test the results).
Bayesian haplotype-based polymorphism discovery and genotyping.
Descriptor library containing a variety of fingerprinting techniques, including the Smooth Overlap of Atomic Positions (SOAP).
A C++ library for parsing and manipulating VCF files.
maeparser is a parser for Schrodinger Maestro files.
A polymorphic bayesian genotyping model with wide applicability.
Open-source, platform-independent, community-supported software for describing and comparing microbial communities
Structural variant and indel caller for mapped sequencing data.
Tandem repeat genotyping with long reads, being a modified version of HipSTR.
Collection of tools for working with BAM files.
VCF manipulation and statistics (e.g. linkage disequilibrium, allele frequency, Fst).
A system for rapidly aligning entire genomes, whether in complete or draft form.
SKESA is a de-novo sequence read assembler for microbial genomes. It uses conservative heuristics and is designed to create breaks at repeat regions in the genome. This leads to excellent sequence quality without significantly compromising contiguity.
Scalable gVCF merging and joint variant calling for population sequencing projects
A generic C++ trie search tree library for small alphabets, allowing customizable leaf node structures and supporting approximate matching and word generation.
A suite of algorithms for matching position weight matrices (PWM) against DNA sequences. It features advanced matrix matching algorithms implemented in C++ that can be used to scan hundreds of matrices against chromosome-sized sequences in few seconds. MOODS can also process high-order PWMs with dependencies between adjacent positions and sequence variants such as SNPs, insertions and deletions.
VerityMap is a tool for mapping long reads to assemblies of extra-long tandem repeats, producing SAM files and identifying potential heterozygous sites and assembly errors through analysis of rare k-mers. It supports PacBio HiFi and ONT reads and generates interactive HTML plots for variant analysis.
Cufflinks assembles transcripts, estimates their abundances, and tests for differential expression and regulation in RNA-Seq samples.
Telseq is a tool for estimating telomere length from whole genome sequence data.
A bioinformatics tool for viewing and calculating base modification frequencies from BAM files
MailsDaddy MBOX to PST Converter is designed to move single or multiple MBOX file to Outlook PST, EML, MSG, HTML etc.