Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

29 of 6,358 resources

Active6132 weeks ago
Shell
MIT

This tool estimates the completeness of KEGG pathway modules from the presence or absence of KEGG orthologues (KOs)

Active403 weeks ago
Shell
Apache-2.0

RAiSD-AI is a tool for training, testing, and deploying Convolutional Neural Networks to detect selective sweeps in genomic data, extending the functionality of the original RAiSD software with machine learning capabilities. It supports SNP data processing, CNN model training with TensorFlow or PyTorch, and genome-wide selective sweep detection.

Active71 month ago
Shell
Other

Databank of optimised macromolecular structures. PDB-REDO entries are refined, rebuilt and validated with one consistent protocol using the equivalent entry in the Protein Data Bank and its experimental data. PDB-REDO entries typically have higher structural quality and a better fit to the experimental data.

Active01 month ago
Shell
BSD-2-Clause

linkset-automation is a set of tools to automatically generates CyTargetLinker linksets from different resources, starting with WikiPathways.

Active01 month ago
Shell
Apache-2.0

Generates pre-miRNA and mature miRNA count tables from read alignments to pre-miRNA sequences and a gff file, both downloaded from mirBase. Produces also read coverage plots of pre-miRNAs.

Active132 months ago
Shell
MIT

blue-crab is a tool to convert from ONT POD5 format to the community maintained SLOW5/BLOW5 format. Lossless nanopore pod5 s/blow5 file conversion.

Active482 months ago
Shell
MIT

SMBGC Annotation using Neural Networks Trained on Interpro Signatures

Active304 months ago
Shell
Apache-2.0

RFantibody is a pipeline for structure-based de novo antibody and nanobody design, integrating backbone design with RFdiffusion, sequence design with ProteinMPNN, and structure prediction with RoseTTAFold2. It provides a comprehensive toolset for generating and filtering high-quality antibody designs.

Active5034 months ago
Shell
MIT

CebraEM is a bioinformatics tool for analyzing and processing large-scale imaging data, providing a pipeline for segmentation, annotation, and analysis with support for both Linux and Windows environments. It includes modules for core functionality, annotation, and network analysis, requiring specific dependencies and a conda environment for execution.

Idle96 months ago
Shell
MIT

Efficient foundation model and benchmark for multi-species genome understanding with context-aware nucleotide representations, improving upon DNABERT for diverse genomic task transfer learning (UIUC MAGICS Lab, 484+ stars)

Idle5016 months ago
Shell
Apache-2.0

The Data Science Ontology is a research project of IBM Research AI and Stanford University Statistics. Its long-term objective is to improve the efficiency and transparency of collaborative, data-driven science.

Idle4210 months ago
Shell
CC-BY-4.0

The midlevel energy ontology (MENO) is a BFO-based midlevel ontology. It comprises the concepts for energy qualities, energy-based dispositions and energy-driven transformation and transfer processes and their interrelations. It has the goal to provide an upper level structure for these concepts for energy-related domain ontologies.

Idle21 year ago
Shell
Idle21 year ago
Shell
Idle121 year ago
Shell
CC-BY-4.0
Idle41 year ago
Shell

Short Python script (using Biopython library functions) to extract sequences from a FASTA, QUAL, FASTQ, or SFF file based on the list of IDs given by a column of a tabular file. The output order follows that of the tabular file, and if there are duplicates in the tabular file, there will be duplicates in the output sequence file.

Stale172 years ago
Shell
Other
Archived762 years ago
Shell
Apache-2.0

Syntax Highlighting for Computational Biology file formats (SAM, VCF, GTF, FASTA, PDB, etc...) in vim/less/gedit/sublime.

Stale2733 years ago
Shell
GPL-3.0

SEACR: Sparse Enrichment Analysis for CUT&RUN

Stale1223 years ago
Shell
GPL-2.0
Stale54 years ago
Shell
CC0-1.0

Subvolume processing scripts with the TOM toolbox is a collection of scripts form a pipeline for subvolume alignment and averaging of electron cryo-tomography data.

Stale94 years ago
Shell
MIT

A cross-system scripting language for working with big data pipelines in computer systems of different sizes and capabilities.

Stale925 years ago
Shell

thromboSeq is a bioinformatics tool designed for the analysis of thrombosis-related sequencing data, providing functionalities for variant calling, annotation, and functional interpretation. It streamlines the processing of high-throughput sequencing data to identify genetic variants associated with thrombotic disorders.

Stale47 years ago
Shell
GPL-3.0

Virtual machine with all software and sample data to run 3D-e-Chem Knime workflows

Stale177 years ago
Shell
Apache-2.0

RxDock is a fast and versatile open-source docking program that can be used to dock small molecules against proteins and nucleic acids. It is designed for high-throughput virtual screening (HTVS) campaigns and binding mode prediction studies.

Bin Chicken - recovery of low abundance and taxonomically targeted metagenome assembled genomes (MAGs) through strategic coassembly

Simple test framework for Nextflow pipelines