Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source(1)
Type
126 of 6,573 resources
Showing 51–100
GFF and GTF file manipulation and interconversion.
A C++ library for parsing and manipulating VCF files.
Deep learning-based variant caller
Genetic variant annotation and effect prediction toolbox.
FASTQ and SAM quality control using Python.
BWA-MEM drop-in replacement: 2-3x faster, 2-5x cheaper, 100% identical output on standard CPUs.
lumpy: a general probabilistic framework for structural variant discovery.
A polymorphic bayesian genotyping model with wide applicability.
A generic but comprehensive bacterial annotation pipeline, built with Nextflow, with nice graphical options for investigating results.
a specification for describing analysis workflows and tools that are portable and scalable across a variety of software and hardware environments, from workstations to cluster, cloud, and high performance computing (HPC) environments.
Prokka: rapid prokaryotic genome annotation. Prokka is one of the most cited annotation command line tools for microbial genome annotations.
Biocaml aims to be a high-performance user-friendly library for Bioinformatics.
Sort genomic files according to a specified order.
Structural variant and indel caller for mapped sequencing data.
SIMD C library for global, semi-global, and local pairwise sequence alignments
Easily get SRA download links and other information.
Toolkit for processing sequences in FASTA/Q formats.
GRIDSS: the Genomic Rearrangement IDentification Software Suite.
Collection of tools for working with BAM files.
VCF manipulation and statistics (e.g. linkage disequilibrium, allele frequency, Fst).
Python wrapper for [bedtools](https://github.com/arq5x/bedtools).
A system for rapidly aligning entire genomes, whether in complete or draft form.
SKESA is a de-novo sequence read assembler for microbial genomes. It uses conservative heuristics and is designed to create breaks at repeat regions in the genome. This leads to excellent sequence quality without significantly compromising contiguity.
Batteries included genomic analysis pipeline for variant and RNA-Seq analysis, structural variant calling, annotation, and prediction.
Workflow library embedded in the Go programming language, focusing on supporting complex workflow constructs, compiling to a single binary, providing powerful file naming and comprehensive audit reports for every output
Resources on ChIP-seq data which include papers, methods, links to software, and analysis.
A pipeline for preprocessing short and long sequencing reads, built with Nextflow.
Partial-Order Alignment for fast alignment and consensus of multiple homologous sequences.
structural variant calling and genotyping with existing tools, but,smoothly.
A collection of research papers for AI-based protein design.
Scalable gVCF merging and joint variant calling for population sequencing projects
Solid path for those of you who want to complete a Bioinformatics course on your own time, for free, with courses from the best universities in the World.
file format conversion in Biopython in a convenient way.
Predicts whether an amino acid substitution affects protein function.
Write-once-read-many table for large datasets.
A fuzzy Bruijn graph approach to long noisy reads assembly
Git repo of useful single line commands.
Displaying sequence statistics for next-generation sequencing.
Educational resource on performing RNA-seq analysis in the cloud using Amazon AWS cloud services. Topics include preparing the data, preprocessing, differential expression, isoform discovery, data visualization, and interpretation.
Easily submitting PBS jobs with script template. Multiple input files supported.
Syntax Highlighting for Computational Biology file formats (SAM, VCF, GTF, FASTA, PDB, etc...) in vim/less/gedit/sublime.
Create an index on a compressed text file.
Point and click, cross platform suite for analysing and visualizing next-generation sequencing datasets.
Go Get Data; A command line interface for obtaining genomic data.
FASTQ/A short-reads pre-processing tools: Demultiplexing, trimming, clipping, quality filtering, and masking utilities.
[@crazyhottommy](https://github.com/crazyhottommy)'s notes on various steps and considerations when doing RNA-seq analysis.