Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

198 of 6,358 resources

Showing 151198

Microsoft's AI-powered ab initio biomolecular dynamics simulation achieving quantum-mechanical accuracy for proteins with 10,000+ atoms, orders of magnitude faster than DFT using protein fragmentation and ML force fields (Nature 2024)

Idle5761 year ago
Python
MIT

Equivariant graph attention Transformer (ICLR2023)

Idle2831 year ago
Python
MIT

Extension of ProteinMPNN for protein sequence design in the context of small-molecule ligands, metal ions, and nucleic acids, enabling binding site engineering and co-factor redesign (Baker Lab)

Idle5881 year ago
Python
MIT

Geometric deep learning model predicting transcriptional outcomes of novel single- and multi-gene perturbations using gene–gene knowledge graphs, 40% higher precision than prior methods on combinatorial perturbation prediction (Stanford, Nature Biotechnology 2024)

Idle3791 year ago
Python
MIT

Large-scale biomolecular instruction dataset for chemistry/biology LLMs (ICLR2024)

Idle2941 year ago
Python
MIT

Tools for adding mutations to existing `.bam` files, used for testing mutation callers.

Idle2511 year ago
Python
MIT

Batteries included genomic analysis pipeline for variant and RNA-Seq analysis, structural variant calling, annotation, and prediction.

Idle1K1 year ago
Python
MIT

Materials informatics benchmark

Idle2071 year ago
Python
MIT

Structure-aware prefix adaptation for integrating LLMs with knowledge graphs (ACM MM 2024)

Idle2121 year ago
Python
MIT

Resources on ChIP-seq data which include papers, methods, links to software, and analysis.

Idle8531 year ago
Python
MIT

UNIX-style FASTA manipulation tools.

Idle171 year ago
Python
MIT

Biomedical text generation

Stale4.5K2 years ago
Python
MIT

A benchmarking platform for molecular generation models.

Stale9772 years ago
Python
MIT

Open language model for mathematics (7B/34B) trained on Proof-Pile-2, outperforming Minerva at equal scale on MATH benchmark, with tool use and formal theorem proving in Lean without finetuning (EleutherAI, ICLR 2024)

Stale1.1K2 years ago
Python
MIT

A package for benchmarking of models for _de novo_ molecular design.

Stale5262 years ago
Python
MIT

Protein structure prediction from ESM models

Archived4.2K2 years ago
Python
MIT

OpenChem is a deep learning toolkit for Computational Chemistry with PyTorch backend.

Stale7472 years ago
Python
MIT

Secure text-to-visualization through standardized chart specifications

Stale2822 years ago
Python
MIT

First foundation model for weather and climate by Microsoft, Vision Transformer-based architecture trained on heterogeneous datasets (ICML 2023)

Stale7012 years ago
Python
MIT

Screen a bacterial assembly (contigs/CDS or proteins) for nucleotide or protein sequences. Pipeline that screens for presence of genes of interest (GOI) in bacterial assemblies. Generates multiple CSVs and plots that describe which genes are present and how variable their sequence is. Can use DNA or protein query sequences (GOIs) and DNA contigs/fastas or protein fastas as database (db) to search in.

Stale62 years ago
Python
MIT

An open, extensible Python framework for GPU-accelerated alchemical free energy calculations.

Stale2023 years ago
Python
MIT

Easily submitting PBS jobs with script template. Multiple input files supported.

Stale293 years ago
Python
MIT

A Library for Deep Learning in Biology and Chemistry.

Stale7023 years ago
Python
MIT

A deep learning framework (based on Chainer) with applications in Biology and Chemistry.

Stale7023 years ago
Python
MIT

A platform for graph-based molecular generation using graph neural networks.

Archived3813 years ago
Python
MIT

Enables machine learning on three-dimensional molecular structure.

Stale3193 years ago
Python
MIT

a robust molecular representation learning framework against distribution shifts.

Stale613 years ago
Python
MIT

Go Get Data; A command line interface for obtaining genomic data.

Stale423 years ago
Python
MIT

A cookiecutter template for bioinformatics projects, with a focus on building bioinformatics workflows that can run on the MPI-IE cluster according to FAIR principles.

Stale133 years ago
Python
MIT

Hierarchical Generation of Molecular Graphs using Structural Motifs.

Stale4414 years ago
Python
MIT

Spherical CNNs for astronomy

Stale1674 years ago
Python
MIT

[@crazyhottommy](https://github.com/crazyhottommy)'s notes on various steps and considerations when doing RNA-seq analysis.

Stale1.1K4 years ago
Python
MIT

Automated strain separation of low-complexity metagenomes

Stale524 years ago
Python
MIT

Crystal property prediction

Stale8754 years ago
Python
MIT

Computation Pipeline library for python widely used in science and bioinformatics.

Stale1755 years ago
Python
MIT

Easy-to-use DNA sequence visualization tool that turns FASTA files into browser-based visualizations.

Archived425 years ago
Python
MIT

Pythonic access to the UCSC Genome database.

Stale1375 years ago
Python
MIT

NanoSV is a software package that can be used to identify structural genomic variations in long-read sequencing data, such as data produced by Oxford Nanopore Technologies’ MinION, GridION or PromethION instruments, or Pacific Biosciences RSII or Sequel sequencers.

Stale926 years ago
Python
MIT

Automatic Filtering, Trimming, Error Removing and Quality Control for fastq data.

Stale2146 years ago
Python
MIT

Molecule validation and standardization based on [RDKit](http://www.rdkit.org/).

Stale1866 years ago
Python
MIT

A port of [pyVCF](https://github.com/jamescasbon/PyVCF) using Cython for speed.

Stale538 years ago
Python
MIT

Tool to generate a count matrix for expression data in Galaxy. generate_count_matrix reads in one or more input text files with expression counts and produces a single combined file. Each input will have a column in the matrix containing expression values. The column containing gene (or feature) names should be identical for all input count files.

Stale09 years ago
Python
MIT

Membrane Protein-Lipid Interaction Database. A large-scale experimentally validated dataset of 80685 residue-level lipid contact annotations across 4712 membrane proteins derived from PDB crystal and cryo-EM structures. Provides pre-computed binary contact labels, continuous distance values, sequence-identity-based cluster assignments, and ready-made train-validation-test splits for machine learning.

Python package for biodatafuse project.

AmsterdamUMCdb is a database of de-identified health data related to tens of thousands of intensive care unit admissions, including demographics, vital signs, laboratory tests and medications.

Reactr is an modularized, Snakemake workflow for automated, species-agnostic characterization of gene families from sequence to experimental design. Given a query protein sequence and NCBI taxonomy IDs (or RefSeq assembly accessions), reactr retrieves genomic data and runs comprehensive analysis across 4 integrated tiers: (1) evolutionary analysis, including homolog detection, domain-based clustering, multiple sequence alignment, and phylogenetic inference; (2) synteny and selection analysis, detecting collinear blocks and calculating Ka/Ks ratios; (3) structural and regulatory characterization, including motif discovery, chromosomal mapping, biochemical property prediction, subcellular localization prediction, and promoter analysis; and (4) experimental design tools, generating PCR primers and scored CRISPR gRNAs for lab validation. Reactr bridges computational prediction and experimental validation, thus enabling rapid transition from genomic discovery to functional studies.