Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

3,476 of 6,569 resources

Showing 3,4513,476

austraits is an R package for accessing the AusTraits Plant database and working with traits.build databases

Browser-based tool to open almost any file that carries sequence — FASTA, FASTQ, GenBank, EMBL, Swiss-Prot, AB1/ABIF, SCF, Clustal, Stockholm, PHYLIP, NEXUS, MSF, PIR, MEGA, GFF3, SAM, BAM, GFA, PDB and ACE — see every sequence inside, pick the ones you want, and save them as FASTA. The format is detected from the file content, not from the extension, so unlabelled or misnamed files still open, and gzip-compressed files are unpacked in place. Records can be filtered by length, name, GC or sequence type, reverse-complemented, transcribed DNA↔RNA or translated to protein, deduplicated and reordered before saving. Runs entirely in the browser — files are never uploaded.

VigyanLLM is a sovereign, on-premises biomedical AI platform designed for computational biology. It provides autonomous primer design, evaluating melting temperature (Tm) and GC content; CRISPR guide RNA analysis with off-target scoring; BLAST sequence similarity searching; multiple sequence alignment (MSA); and GPU-accelerated molecular docking for drug discovery. Unlike cloud-based SaaS, VigyanLLM deploys via Docker to ensure institutional genomic data sovereignty

AI Agent for Biomedical Research

SciAgentKit is an MCP-native toolkit that connects AI agents to reproducible computational drug-discovery workflows. It integrates established tools for molecular analysis, protein-structure assessment, binding-site detection, molecular docking, molecular dynamics, trajectory analysis and scientific reporting.

Trace4MedicalImageCleaning™ is a tool aimed at automatically detecting and removing text in medical images, with a specific focus on 2D ultrasound and mammography studies.

NIM Studio is a local-first platform for neuroinformatics, BIDS organization, metadata generation, duplicate auditing, and scalable research data management.

A bioinformatics tool for viewing and calculating base modification frequencies from BAM files

Simple test framework for Nextflow pipelines

Reactr is an modularized, Snakemake workflow for automated, species-agnostic characterization of gene families from sequence to experimental design. Given a query protein sequence and NCBI taxonomy IDs (or RefSeq assembly accessions), reactr retrieves genomic data and runs comprehensive analysis across 4 integrated tiers: (1) evolutionary analysis, including homolog detection, domain-based clustering, multiple sequence alignment, and phylogenetic inference; (2) synteny and selection analysis, detecting collinear blocks and calculating Ka/Ks ratios; (3) structural and regulatory characterization, including motif discovery, chromosomal mapping, biochemical property prediction, subcellular localization prediction, and promoter analysis; and (4) experimental design tools, generating PCR primers and scored CRISPR gRNAs for lab validation. Reactr bridges computational prediction and experimental validation, thus enabling rapid transition from genomic discovery to functional studies.

Harmonize numerical values extracted from medical images (e.g. acquired with different models of image-acquisition system)

Browser-based viewer for GenBank and GenPept records — .gb, .gbk, .gbff, .gp and plain GenBank text. Renders an interactive linear and circular feature map, including circular plasmid maps, alongside the annotated source text and the nucleotide/protein sequence. Translates CDS features using the record's own genetic code and translation qualifiers, flags where the stored /translation disagrees with a plain translation, and adds optional computed layers: ORF prediction and restriction-site mapping. Drag a range in the sequence band to select it, then copy that stretch — either strand, as DNA or as protein, plain or FASTA. Filters features by type, handles multi-record files, and keeps open records as local sessions, so a closed tab can be picked up where it was left. Runs entirely in the browser — files are never uploaded.

ArrayAnalysis is a web-based application for transcriptomic data analysis. It supports the analysis of both microarray and RNA-seq data. The tool may also be installed locally as a desktop app, Docker image, or R package.

Supernova is a software package for de novo assembly from Chromium Linked-Reads that are made from a single whole-genome library from an individual DNA source. A key feature of Supernova is that it creates diploid assemblies, thus separately representing maternal and paternal chromosomes over very long distances. Almost all other methods instead merge homologous chromosomes into single incorrect 'consensus' sequences. Supernova is the only practical method for creating diploid assemblies of large genomes.

edf2csv is a local command-line tool for converting EDF, EDF+, BDF, and BDF+ physiological recordings into CSV and JSON files. It exports signal values, channel information, annotations, and recording metadata while preserving original sampling rates, physical units, and discontinuities.

Plotwright is a browser-based statistics and scientific graphing workbench for life-science researchers. It provides a Data to Analysis to Graph workflow for statistical tests, regression and curve fitting, survival and ROC analysis, and editable scientific figures. It can import supported content from modern GraphPad Prism .prism files and reports preserved, approximated, and unsupported objects for review. Projects can be saved as local files. Plotwright is commercial software with a 14-day no-card trial.

Predicting the effect of mutations on protein-RNA binding with Deep Learning | This repository contains all DeepCLIP Python code | A context-aware neural network for modeling and predicting protein binding to nucleic acids using only sequence input | DeepCLIP is a neural network with shallow convolutional layers connected to a bidirectional LSTM layer

Plant Compound Extractor is a desktop application that builds a ready-to-use, deduplicated library of 3D ligand structures for a given plant. It queries multiple natural-product and chemical databases (COCONUT, LOTUS, Wikidata, PubChem, PlantaeDB, USDA Dr. Duke's, KNApSAcK and IMPPAT) in parallel, resolves each compound against PubChem for a canonical structure, and falls back to direct source retrieval when needed. Retrieved structures are then converted to 3D using RDKit, with configurable conformer generation and physicochemical filters (molecular weight, rotatable bonds, ring size, etc.). It can also process a manually supplied compound list, or convert an existing folder of 2D structures to 3D.

Desktop viewer for microscopy and whole slide pathology images on Windows and macOS. Opens whole slide scanner formats (Aperio SVS, Hamamatsu NDPI, MIRAX MRXS, Leica SCN, Ventana BIF) alongside acquisition formats (Zeiss CZI, Nikon ND2, DICOM, TIFF/OME-TIFF) in a single application, providing pyramid navigation of gigapixel images, metadata inspection, calibrated measurements, annotations, and on-device segmentation and object counting.

Pan.bio is a cloud genomics platform for pipeline execution, exploratory analysis, and clinical variant interpretation. Workflows runs validated Nextflow and nf-core pipelines including Sarek, rnaseq, scrnaseq, mag, ampliseq, chipseq and atacseq without local installation. Notebooks provides Python and R sessions with a preinstalled bioinformatics stack, importing public data from GEO, SRA and IPG by accession and reading Workflows outputs directly. VAIC applies ACMG/AMP variant classification with gene-specific rule sets from CanVIG-UK and ClinGen ENIGMA, with automated evidence criteria implemented for BRCA1 and BRCA2. Cohorts provides federated analysis of patient data within a Trusted Research Environment.

Novel machine learning-based method for prediction of ligand binding sites from protein structure.

PhonaLab is a browser-based platform for acoustic analysis of voice recordings, aimed at speech-language pathologists, voice clinicians, and researchers. It computes validated multiparametric acoustic indices — including the Acoustic Voice Quality Index (AVQI), Acoustic Breathiness Index (ABI), smoothed cepstral peak prominence (CPPS), and glottal-to-noise excitation ratio (GNE) — from sustained-vowel and connected-speech recordings, using Praat algorithms via the Parselmouth interface. Audio is processed in memory and not stored. Interface available in English, Brazilian Portuguese, and Spanish.

Gelyze is a web-based image analysis tool for agarose gel electrophoresis of nucleic acids. A user uploads a gel image, confirms the proposed lane positions and the ladder baseline, and a deterministic algorithm then estimates fragment size from ladder calibration and quantifies relative band intensity by densitometry. Measurement and interpretation are kept separate: the numbers come from the algorithm, and an assistive AI layer explains the resulting pattern in plain language without producing any of the values. For research use only.

MailsDaddy MBOX to PST Converter is designed to move single or multiple MBOX file to Outlook PST, EML, MSG, HTML etc.

FT-ITC Analysis is free, open-source desktop software for processing, fitting, and analysing isothermal titration calorimetry (ITC) data. It imports raw data from MicroCal, TA Instruments/NanoAnalyze, and PEAQ ITC project files, as well as integrated heat data and FT-ITC project files. The software supports baseline correction, injection integration, interaction-model fitting, global analysis, uncertainty estimation, advanced thermodynamic analyses, and publication-figure export on macOS, Windows, and Linux.

Browser-based viewer that maps sequencing reads onto one short reference — an amplicon, gene or plasmid. Reads open as Sanger AB1/ABIF, SCF, FASTA, FASTQ or a SAM somebody else already mapped (gzipped files are unpacked in place); the reference as FASTA, GenBank or a read. Both read orientations are tried automatically. The pileup reports per-position depth, where reads disagree with the target, and the consensus — phred-weighted for capillary reads, which keep their chromatogram under the letters. An optional protein lane translates target and consensus side by side. Reads are placed by minimap2 compiled to WebAssembly, or by the built-in aligner. One read or the whole alignment saves as FASTA. Runs entirely in the browser — files are never uploaded.