Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

41 of 7,055 resources

SPAdes (St. Petersburg genome assembler) is an assembly toolkit containing various assembly pipelines and the de-facto standard for prokaryotic genome assemblies.

Active9633 days ago
C++
NOASSERTION

Universal molecular toolkit that can be used for molecular fingerprinting, substructure search, and molecular visualization written in C++ package, with Java, C#, and Python wrappers.

Active4081 week ago
C++
Apache-2.0

An ultrafast protein aligner for `blastp` and `blastx` like searches.

Active1.3K1 week ago
C++
GPL-3.0

High-performance molecular simulation toolkit

Active2K2 weeks ago
C++

An ultrafast and memory-efficient tool for aligning sequencing reads to long reference sequences.

Active8113 weeks ago
C++
GPL-3.0

Bayesian haplotype-based polymorphism discovery and genotyping.

Active8811 month ago
C++
MIT

The modern C++ library for sequence analysis.

Active4611 month ago
C++
NOASSERTION

Oxford Nanopore's official deep-learning basecaller for nanopore sequencing, converting raw electrical signals into DNA/RNA sequences with integrated modified-base (methylation) detection and efficient CPU/GPU inference; foundational tool for long-read genomics, epigenetics, and real-time sequencing analysis (nanoporetech, 846+ stars, actively maintained)

Active8661 month ago
C++
NOASSERTION

Structural variant discovery by integrated paired-end and split-read analysis.

Active5341 month ago
C++
BSD-3-Clause

LLMs as copilots for theorem proving in Lean 4, exposing native tactics (`suggest_tactics`, `search_proof`, `select_premises`) that embed language model inference and premise retrieval directly inside the Lean proof environment, supporting local CTranslate2/CUDA inference as well as remote model APIs for interactive and automated proof search (Caltech & NVIDIA, NeurIPS 2024, 1.2K+ stars)

Active1.3K1 month ago
C++
MIT

Open-source, platform-independent, community-supported software for describing and comparing microbial communities

Active2801 month ago
C++
GPL-3.0

A collection of object-oriented software tools for problems involving chemical kinetics, thermodynamics, and transport processes.

Active8381 month ago
C++
NOASSERTION

A software package for estimating gene and isoform expression levels from RNA-Seq data.

Active4792 months ago
C++
GPL-3.0

Genome mapping and spliced alignment of cDNA or amino acid sequences

Active1143 months ago
C++
GPL-2.0

Deep learning framework for molecular docking extending AutoDock Vina with convolutional neural network scoring functions, achieving superior virtual screening enrichment and pose prediction across diverse target classes; widely adopted in pharmaceutical structure-based drug design (J. Cheminformatics, 915+ stars, actively maintained)

Active9733 months ago
C++
Apache-2.0

A single molecule sequence assembler for genomes large and small.

Active7083 months ago
C++

A haplotype-resolved assembler for accurate Hifi reads.

Active8064 months ago
C++
MIT

A small <720Kb C++ windows utility. That allows you to load Ancestry, 23andMe, FTDNA, or Genes for Good RAW DNA files search them, merge them. covert them to Ancestry format. But also create files from peer reviewed publications to compare with you loaded data to give your genetic disposition for the condition you have entered the data for an statistical risk if OR values are included. Included with the program are example files for Type 2 Diabetes risk factors. (As I have type 2 Diabetes so I could test the results).

Active04 months ago
C++
GPL-3.0

Descriptor library containing a variety of fingerprinting techniques, including the Smooth Overlap of Atomic Positions (SOAP).

Active4755 months ago
C++
Apache-2.0

A C++ library for parsing and manipulating VCF files.

Idle6866 months ago
C++
MIT

maeparser is a parser for Schrodinger Maestro files.

Idle297 months ago
C++
MIT

A polymorphic bayesian genotyping model with wide applicability.

Idle3267 months ago
C++
MIT

Structural variant and indel caller for mapped sequencing data.

Archived46712 months ago
C++
NOASSERTION

Tandem repeat genotyping with long reads, being a modified version of HipSTR.

Idle401 year ago
C++
GPL-2.0

Collection of tools for working with BAM files.

Idle4321 year ago
C++
MIT

VCF manipulation and statistics (e.g. linkage disequilibrium, allele frequency, Fst).

Idle5631 year ago
C++
LGPL-3.0

A system for rapidly aligning entire genomes, whether in complete or draft form.

Idle5751 year ago
C++
Artistic-2.0

SKESA is a de-novo sequence read assembler for microbial genomes. It uses conservative heuristics and is designed to create breaks at repeat regions in the genome. This leads to excellent sequence quality without significantly compromising contiguity.

Idle1271 year ago
C++
NOASSERTION

adapter trimmer for Oxford Nanopore reads

Stale3872 years ago
C++
GPL-3.0

Scalable gVCF merging and joint variant calling for population sequencing projects

Stale1892 years ago
C++
Apache-2.0

A generic C++ trie search tree library for small alphabets, allowing customizable leaf node structures and supporting approximate matching and word generation.

Stale12 years ago
C++
MIT

A suite of algorithms for matching position weight matrices (PWM) against DNA sequences. It features advanced matrix matching algorithms implemented in C++ that can be used to scan hundreds of matrices against chromosome-sized sequences in few seconds. MOODS can also process high-order PWMs with dependencies between adjacent positions and sequence variants such as SNPs, insertions and deletions.

Stale1183 years ago
C++
NOASSERTION

VerityMap is a tool for mapping long reads to assemblies of extra-long tandem repeats, producing SAM files and identifying potential heterozygous sites and assembly errors through analysis of rare k-mers. It supports PacBio HiFi and ONT reads and generates interactive HTML plots for variant analysis.

Stale393 years ago
C++
GPL-3.0

Cufflinks assembles transcripts, estimates their abundances, and tests for differential expression and regulation in RNA-Seq samples.

Stale3256 years ago
C++
BSL-1.0

Telseq is a tool for estimating telomere length from whole genome sequence data.

Stale787 years ago
C++
GPL-3.0

The fmcsR package introduces an efficient maximum common substructure (MCS) algorithms combined with a novel matching strategy that allows for atom and/or bond mismatches in the substructures shared among two small molecules. The resulting flexible MCSs (FMCSs) are often larger than strict MCSs, resulting in the identification of more common features in their source structures, as well as a higher sensitivity in finding compounds with weak structural similarities. The fmcsR package provides several utilities to use the FMCS algorithm for pairwise compound comparisons, structure similarity searching and clustering.

Stale610 years ago
C++

A bioinformatics tool for viewing and calculating base modification frequencies from BAM files

MailsDaddy MBOX to PST Converter is designed to move single or multiple MBOX file to Outlook PST, EML, MSG, HTML etc.

This package provides an R wrapper of the popular bowtie2 sequencing reads aligner and AdapterRemoval, a convenient tool for rapid adapter trimming, identification, and read merging. The package contains wrapper functions that allow for genome indexing and alignment to those indexes. The package also allows for the creation of .bam files via Rsamtools.

This package provides a framework for the quantification and analysis of Short Reads. It covers a complete workflow starting from raw sequence reads, over creation of alignments and quality control plots, to the quantification of genomic regions of interest. Read alignments are either generated through Rbowtie (data from DNA/ChIP/ATAC/Bis-seq experiments) or Rhisat2 (data from RNA-seq experiments that require spliced alignments), or can be provided in the form of bam files.

Provides a high-level R interface to CoreArray Genomic Data Structure (GDS) data files. GDS is portable across platforms with hierarchical structure to store multiple scalable array-oriented data sets with metadata information. It is suited for large-scale datasets, especially for data which are much larger than the available random-access memory. The gdsfmt package offers the efficient operations specifically designed for integers of less than 8 bits, since a diploid genotype, like single-nucleotide polymorphism (SNP), usually occupies fewer bits than a byte. Data compression and decompression are available with relatively efficient random access. It is also allowed to read a GDS file in parallel with multiple R processes supported by the package parallel.