Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain(1)
Language
License
Source
Type
17 of 7,064 resources
AcinetoScope is an automated, comprehensive bioinformatics pipeline designed specifically for the genomic analysis of Acinetobacter baumannii, a WHO Critical Priority pathogen responsible for devastating hospital-acquired infections. It integrates seven analysis types (MLST, ABRicate, AMRFinder, Kaptive 3, APT, PlasmidFinder, and mutation detection) into a single automated workflow — from FASTA to actionable insights. The pipeline offers both gene-centric and sample-centric reporting, dynamic grouping by typing, and is optimised for HPC, cloud, and container environments.
Module for single-cell data extraction given a segmentation mask and multi-channel image.
HIDE-Deconv is a framework for characterizing cellular remodeling from bulk RNA-seq data using hierarchical cell-type deconvolution across multiple levels of cellular resolution. It provides an integrated workflow for single-cell reference preprocessing, estimation of cellular compositions, and downstream analysis of deconvolution results, including visualization, clustering, differential composition, and survival analysis.
A python extension, written in C, for quick access to bigBed files and access to and creation of bigWig files.
A comprehensive R package for identifying and ranking influential nodes in biological and other complex networks. The package implements the Integrated Value of Influence (IVI), Experimental data-based Integrative Ranking (ExIR), SIRIR, and numerous network centrality measures, enabling network topology analysis, influential node detection, feature prioritization, and candidate biomarker discovery. It also provides functions for network reconstruction, centrality assessment, visualization, and analysis of relationships between centrality measures.
MiRA (Multilayer Interactive Rendering Application) is a free, browser-based, installation-free tool for interactive visualisation of multilayer networks in ecology and biology. While developed with biology in mind, MiRA can render any multilayer network. Available as a standalone web app and also fully integrated into the emln R package.
SQUARNA is a tool for RNA secondary structure prediction. It can take a single RNA sequence or an alignment of sequences as input. SQUARNA handles pseudoknots and can predict alternative structures. SQUARNA allows structural restraints and chemical probing data as additional input and is available at https://github.com/febos/SQUARNA and https://larnal.imol.institute/.
Documentation Rectangle is an open-source Python package for single-cell-informed cell-type deconvolution of bulk and spatial transcriptomic data. Rectangle presents a novel approach to second-generation deconvolution, characterized by hierarchical signature building for fine-grained cell-type deconvolution, estimation and correction of unknown cellular content, and efficient handling of large-scale single-cell data during signature matrix computation. Rectangle was developed to overcome the current challenges in cell-type deconvolution, providing a robust and accurate methodology while ensuring a low computational profile.
Toolbox for comparative genomics of MAGs
SMBGC Annotation using Neural Networks Trained on Interpro Signatures
DeepConsensus uses gap-aware sequence transformers to correct errors in Pacific Biosciences (PacBio) Circular Consensus Sequencing (CCS) data.
A database system designed to store, organize, and manage large-scale nucleotide sequencing read data (like PacBio reads) for the Dazzler genome assembler
Utility that performs integrated analyses of 'gene' data (a set of genes or other genomic features) with 'peak' data (a set of regions, for example ChIP peaks) to identify the genes nearest to each peak, and vice versa.
Short Python script (using Biopython library functions) to extract sequences from a FASTA, QUAL, FASTQ, or SFF file based on the list of IDs given by a column of a tabular file. The output order follows that of the tabular file, and if there are duplicates in the tabular file, there will be duplicates in the output sequence file.
GFF3sort: A Perl Script to sort gff3 files and produce suitable results for tabix tools
NanoporeDB is an open-access structural database dedicated to the exploration, analysis, and design of protein nanopores, which serve as essential molecular gateways in biological membranes and form the basis of many advanced biosensing and sequencing technologies. This platform integrates large-scale structure-guided mining and deep learning-based modeling using AlphaFold-Multimer and AlphaFold3 to provide about 7,000 high-confidence multimeric nanopore structures. Each entry includes detailed information on membrane embedding, pore geometry annotation, and constriction profiling to support functional and biophysical interpretation. Through an interactive 3D visualization interface and quantitative parameters such as tilt angle, insertion depth, and pore geometry, NanoporeDB enables researchers to explore nanopore diversity, discover novel scaffolds, and accelerate innovation in molecular sensing, precision diagnostics, and synthetic biology.
Thoa is a cloud bioinformatics platform. Write your Nextflow or Snakemake pipeline, point it at your data, and Thoa handles the rest: provisioning VMs (up to 12TB RAM), resolving dependencies, managing execution. No cloud expertise needed. Every job captures its full context:data, software versions, environment, machine specs, as a reproducibility artifact. Share it with a colleague and they can view or re-run the analysis without an account. Key features: AI debugger that fixes environment and dependency issues in real time. Pipeline tracking with per-step telemetry. if step 47 of 200 fails, re-run from there, not from scratch. One-click data sharing without registration. AI-assisted workflow creation from plain English. Free tier available. Starter $35/mo, Pro $109/mo, Team $480/mo. Zero-egress storage. Based in Zug, Switzerland. thoa.io