Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

23 of 7,050 resources

dadi is a bioinformatics tool for inferring demographic history and selection from genetic data using diffusion approximations, offering speed and flexibility in modeling population dynamics. It supports up to three populations with customizable parameters and provides efficient computational performance.

Active82 weeks ago
Python
NOASSERTION

Offline tool for cleaning tables of human gene and protein identifiers (TXT, CSV, TSV, XLSX). It maps approved symbols, aliases, previous symbols, Ensembl gene, UniProt, Entrez, RefSeq and HGNC identifiers to current HGNC approved symbols with cross-references, using a bundled HGNC snapshot. Excel date-corrupted symbols are recovered where the original is unambiguous and flagged for manual review otherwise; no input row is dropped. Each run records the tool version and HGNC release and writes a per-row audit table.

Active03 weeks ago
Python
MIT

CLIFinder is a Galaxy tool designed to identify potential L1 Chimeric Transcripts from RNA-seq data by analyzing paired-end reads in the human genome. It allows customization to detect transcripts initiated by different repeat elements.

Active33 weeks ago
Perl
GPL-3.0

BacSelect is a reproducible resource for selecting compact, nested panels of complete bacterial genome assemblies that span genome architecture. From a defined and versioned public genome universe, BacSelect deterministically selects assemblies so users can choose a panel size from N=10–500 while retaining nestedness across scales. Versioned releases provide genome-panel manifests, downloadable assemblies, structural-coverage information, checksums and provenance for reproducible benchmarking, method development and comparative genomics.

Active04 weeks ago
Python
MIT

Eukaryotic Genome Annotation Pipeline-External caller scripts and documentation

Active2071 month ago
Nextflow
NOASSERTION

Workflow optimized for the analysis of rare diseases, designed to detect SNVs, INDELs , CNVs and SVs in targeted sequencing data (CES/WES) and whole genome sequencing (WGS), built on Nextflow and following nf-core standards. It has an advanced variant annotation optimized for rare diseases diagnosis and discovery.

Active11 month ago
Nextflow
MIT

GAIn is a platform for annotating genetic variants, genomic positions, and regions with reproducible, declarative pipelines using curated Genomic Resource Repositories.

Active11 month ago
Python
MIT

blue-crab is a tool to convert from ONT POD5 format to the community maintained SLOW5/BLOW5 format. Lossless nanopore pod5 s/blow5 file conversion.

Active491 month ago
Python
MIT

RAiSD-AI is a tool for training, testing, and deploying Convolutional Neural Networks to detect selective sweeps in genomic data, extending the functionality of the original RAiSD software with machine learning capabilities. It supports SNP data processing, CNN model training with TensorFlow or PyTorch, and genome-wide selective sweep detection.

Active81 month ago
Python

Exact, validated excision of coordinate-defined genomic regions from transposed NEXUS matrices.

Active11 month ago
Python
MIT

Rapid & standardized annotation of bacterial genomes, MAGs & plasmids

Active6802 months ago
Python
GPL-3.0

A Python script that converts positional information from a SAM dataset into interval format with 0-based start and 1-based end. CIGAR string of SAM format is used to compute the end coordinate.

Active372 months ago
Python
NOASSERTION

Pathogensurveillance is a population genomics pipeline for pathogen identification, variant detection, and biosurveillance. The pipeline accepts paths to raw reads for one or more organisms and creates reports in the form of an interactive HTML document. Significant features include the ability to analyze unidentified eukaryotic and prokaryotic samples, creation of reports for multiple user-defined groupings of samples, automated discovery and downloading of reference assemblies from NCBI RefSeq, and rapid initial identification based on k-mer sketches followed by a more robust multi gene phylogeny and SNP-based phylogeny.

Active612 months ago
Nextflow
MIT

D-GENIES – for Dot plot large Genomes in an Interactive, Efficient and Simple way – is an online tool designed to compare two genomes. It supports large genome and you can interact with the dot plot to improve the visualization.

Active1164 months ago
Shell
GPL-3.0

SCENIC+ is a python package to build gene regulatory networks (GRNs) using combined or separate single-cell gene expression (scRNA-seq) and single-cell chromatin accessibility (scATAC-seq) data.

Idle2638 months ago
Jupyter Notebook
NOASSERTION

Pairwise SNP distance matrix from a FASTA sequence alignment

Idle1569 months ago
C
GPL-3.0

Minigraph is a sequence-to-graph mapper and graph constructor. For graph generation, it aligns a query sequence against a sequence graph and incrementally augments an existing graph with long query subsequences diverged from the graph.

Idle4871 year ago
C
MIT

Scalable gVCF merging and joint variant calling for population sequencing projects

Stale1892 years ago
C++
Apache-2.0

NuclearPhaser is a method for phasing of dikaryotic genomes into the two haplotypes using Hi-C contact graphs. This is an overview of the phasing pipeline for dikaryons.

Stale133 years ago
Python
GPL-3.0

Finds SNP sites from a multi-FASTA alignment file.

Stale2795 years ago
C
NOASSERTION

This package does nucleosome positioning using informative Multinomial-Dirichlet prior in a t-mixture with reversible jump estimation of nucleosome positions for genome-wide profiling.

Stale36 years ago
R

Pan.bio is a cloud genomics platform for pipeline execution, exploratory analysis, and clinical variant interpretation. Workflows runs validated Nextflow and nf-core pipelines including Sarek, rnaseq, scrnaseq, mag, ampliseq, chipseq and atacseq without local installation. Notebooks provides Python and R sessions with a preinstalled bioinformatics stack, importing public data from GEO, SRA and IPG by accession and reading Workflows outputs directly. VAIC applies ACMG/AMP variant classification with gene-specific rule sets from CanVIG-UK and ClinGen ENIGMA, with automated evidence criteria implemented for BRCA1 and BRCA2. Cohorts provides federated analysis of patient data within a Trusted Research Environment.

This package provides a framework for the quantification and analysis of Short Reads. It covers a complete workflow starting from raw sequence reads, over creation of alignments and quality control plots, to the quantification of genomic regions of interest. Read alignments are either generated through Rbowtie (data from DNA/ChIP/ATAC/Bis-seq experiments) or Rhisat2 (data from RNA-seq experiments that require spliced alignments), or can be provided in the form of bam files.