Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language(1)
License
Source
Type
1,191 of 7,068 resources
Showing 351–400
AI-assisted structural engineering workspace for AEC workflows: natural language to structural model, analysis, code-check, and report (171+ stars, MIT License, 2026)
Google DeepMind's diffusion-based ensemble weather forecasting model at 0.25° resolution, outperforming ECMWF ENS on 97.2% of targets up to 15 days ahead, with open-source code and weights (Nature 2024)
Parrotlet-a 2.5 Pro is a purpose-built automatic speech recognition (ASR) model for medical speech in Indian healthcare settings. It transcribes Indian English, Hindi, Marathi, Kannada and Telugu, including the heavily code-mixed speech typical of real consultations (English drug names and clinical…
linkset-automation is a set of tools to automatically generates CyTargetLinker linksets from different resources, starting with WikiPathways.
Neurazum/VLbai-2.6AD
by NeurazumA clinical reasoning assistant for early-stage Alzheimer's assessment. It joins a 3D MRI + biomarker classifier (Vbai-2.6AD) to a reasoning LLM (Gemma 4 12B) inside a single forward pass — the diagnosis is passed as vectors, not text.
Open-source JAX-based software suite for variational optimization of deep-learning molecular wave functions, solving electronic ground and excited states via neural-network trial wave functions with configurable FermiNet, PauliNet, Psiformer, LapNet, and DeepErwin ansätze, geometric transferability across molecular configurations, and effective core potential support (FU Berlin / Noé group, J. Chem. Phys. 2023, 420+ stars, MIT License)
U-Net-style deep neural network for P/S seismic arrival-time picking trained on millions of waveforms from the Northern California Earthquake Data Center, achieving near-analyst picking precision at orders-of-magnitude higher speed and robustness to low signal-to-noise traces where STA/LTA fails; a foundational reference for deep-learning phase picking, integrated into SeisBench model collections and national seismic networks, with PhaseNet-DAS extending it to distributed acoustic sensing (Stanford AI4EPS, 386+ stars, MIT License, actively maintained)
Human-centered research OS with terminal-first harness and local browser Studio, turning research work into reproducible artifact-backed runs through a 9-stage workflow with human approval gates, resume/rollback controls, and venue-aware manuscript packaging (1K+ stars, 2026)
EcoliTyper is a revolutionary bioinformatics pipeline that eliminates workflow fragmentation in E. coli genomic surveillance. By integrating nine core analyses into a single automated workflow, EcoliTyper transforms disconnected genomic data into coherent biological narratives with actionable public health intelligence. It is a species-optimized computational pipeline for comprehensive genotyping and surveillance of Escherichia coli, perfect for clinical microbiology, outbreak investigations, and genomic research.
ProSeqGO predicts Gene Ontology (GO) terms for protein sequences using ESM2 embeddings and a trained 1-Dimensional Convolutional Neural Network multi-label classifier. By integrating recent advances in protein language models, ProSeqGO facilitates large-scale, automated functional annotation directly from sequence input, empowering researchers to infer protein function, explore biological mechanisms, and accelerate discovery in genomics and proteomics.
GrimSqueaker/ProtSent-V2-ESMC-300M
by GrimSqueakerContrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
RiSPICE (Rice SNP Prioritization Integrating Chromatin Effects) is a computational framework for prioritizing non-coding rice variants by integrating predicted chromatin effects from a fine-tuned DNA language model.
Whole-slide pathology foundation model trained on 1.3 billion image tiles from 171K slides using a LongNet-based architecture to encode gigapixel-scale WSIs for cancer subtyping and biomarker prediction (Microsoft Research & Providence, 601+ stars)
prov-gigapath/prov-gigapath-flash
by prov-gigapathprov-gigapath/prov-gigapath
by prov-gigapathVision foundation model for the tree of life, pretrained on diverse biological imagery across taxa for zero-shot species identification, trait extraction, and biodiversity research (Ohio State University Imageomics Institute)
Aignostics/RudolfV-2-S
by AignosticsAignostics/RudolfV-2-B
by AignosticsAignostics/RudolfV-2
by AignosticsHuggingFaceBio/Carbon-3B
by HuggingFaceBioTechnical Report 🧬
Rapid & standardized annotation of bacterial genomes, MAGs & plasmids
ECMWF's unified framework and command-line tool to run AI-based weather forecasting models (GraphCast, Aurora, Pangu, NeuralGCM, FourCastNet) with operational ECMWF data infrastructure, enabling standardized inference and benchmarking across state-of-the-art meteorological AI systems (ECMWF, 576+ stars)
This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…
MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…
mradermacher/Gemma-2B-Uncensored-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
insilicomedicine/Qwen3-0.6B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-0.6B, is the smallest family member and was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-0.6B on aging-related multi-omics and…
Learning operators in Fourier space
Tools for adding mutations to existing `.bam` files, used for testing mutation callers.
!Format !Task !Params !Type !License
Fast, differentiable, JIT-free finite element library for PyTorch enabling GPU-native PDE solving with native autograd, tensorized assembly, and sparse linear algebra; part of the TensorGalerkin framework (218+ stars, Apache 2.0)
llmithull/HealthGPT-LoRA
by llmithullHealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.
Python computational framework for analysis of single-molecule FRET data
GrimSqueaker/ProtSent-V2.5-35M
by GrimSqueakerProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.
Plain-text, git-tracked electronic lab notebook (ELN) for reproducible bioinformatics — threads your R & Python figures into living lab notes with full provenance. Built for single-cell / CyTOF / flow cytometry; works with Obsidian, Quarto & Jupyter.
Neural network-based cryo-EM heterogeneous reconstruction, modeling continuous 3D structure distributions from single-particle images, with CryoDRGN-ET extending to in-cell cryo-electron tomography (MIT CSAIL, Nature Methods 2021/2024)
Utilities for working with CSV/Tab-delimited files.
Probabilistic framework for inferring cell fate decisions and trajectory dynamics from multi-view single-cell data using Markov chains and machine learning, integrating RNA velocity, pseudotime, and metabolic labeling to predict differentiation paths and terminal states (scverse/Theis Lab, 449+ stars, BSD 3-Clause)
Documentation Rectangle is an open-source Python package for single-cell-informed cell-type deconvolution of bulk and spatial transcriptomic data. Rectangle presents a novel approach to second-generation deconvolution, characterized by hierarchical signature building for fine-grained cell-type deconvolution, estimation and correction of unknown cellular content, and efficient handling of large-scale single-cell data during signature matrix computation. Rectangle was developed to overcome the current challenges in cell-type deconvolution, providing a robust and accurate methodology while ensuring a low computational profile.
A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: smooth-L1 regression onto the frozen target, no negatives (the JEPA move).
A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: SigLIP sigmoid pairwise loss.
Computational toolbox for large scale Calcium Imaging Analysis, including movie handling, motion correction, source extraction, spike deconvolution and result visualization, using machine learning for automated neuron detection and activity inference in two-photon and one-photon calcium imaging data (723+ stars, actively maintained)
Interaction Fingerprints for protein-ligand complexes and more.
The Zebrafish Activity Prediction Benchmark for forecasting cellular-resolution neural activity throughout an entire vertebrate brain, combining light-sheet microscopy calcium-imaging data, forecasting tasks, and evaluation tools to advance whole-brain neural dynamics modeling (77+ stars, Apache 2.0)
Next-generation benchmark for data-driven global weather models with standardized evaluation framework and curated datasets for ML forecasting (Google Research, 2024)
Trainable PyTorch reproduction of AlphaFold 3
Simulations of spiking neural networks.
Medical large vision-language model unifying comprehension and generation via heterogeneous knowledge adaptation, enabling holistic medical image understanding, visual question answering, and clinical report generation across diverse modalities (ZJU4HealthCare, 1.6K+ stars)
Meddies/meddies-pii
by MeddiesA multilingual PII extractor for teams that need structured JSON from clinical and administrative text.
GrimSqueaker/ProtSent-V2-150M
by GrimSqueakerContrastively fine-tuned ESM-2 150M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.