Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source
Type(1)
869 of 7,050 resources
Showing 51–100
biohub/ESMC-300M
by biohubESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…
biohub/ESMFold2
by biohubESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…
biohub/ESMFold2-Fast
by biohubESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…
scrc-dnai/DNT-8M-CPL-SMP
by scrc-dnaiscrc-dnai/DNT-100M_PRE-CPL-2e4
by scrc-dnaishreyansh12183/olmo2-7b-biomed-chem
by shreyansh12183A domain-adapted small language model specialized for molecular conformation analysis, computational biochemistry, and pharmacology reasoning.
Aquiles-ai/Evo2-7B
by Aquiles-aiEvo2-7B (Transformers port)
Aquiles-ai/Evo2-1B-Base
by Aquiles-aiEvo2-1B-Base (Transformers port)
Creeperbbs/RMMol
by CreeperbbsThis bundle contains standardized CSV files and a Jupyter sample notebook for RMMol frozen-embedding examples across Biophysics, Physiological, Physical Chemistry, and Quantum Mechanics.
sciai-lab/structures25
by sciai-labMachine-learned orbital-free density functional theory
ufmg-digital-pathology/breast-cancer-virtual-staining
by ufmg-digital-pathologyCycleGAN generators that synthesise Ki-67 and pHH3 immunohistochemistry (IHC) appearance from H&E histopathology tiles of triple-negative breast cancer (TNBC).
!Benchmark card: Hertz 0.7F vs same-size models
ctheodoris/Geneformer
by ctheodoris# Geneformer Geneformer is a foundational transformer model pretrained on a large-scale corpus of human single cell transcriptomes to enable context-aware predictions in settings with limited data in network biology.
tagirshin/RxnScribe-ONNX
by tagirshinONNX conversions of the existing RxnScribe, MolScribe and English EasyOCR checkpoints. This is a community conversion, not a newly trained model or an upstream release. Use the matching RxnScribe feature branch and its RxnScribeONNX interface. PyTorch is needed for export, not inference.
DOEJGI/vhamster-models
by DOEJGInayoung10/Packora-ckpt
by nayoung10- Project page - Try Packora (live demo) - Paper - Code - Dataset manifests - Hugging Face collection
InstaDeepAI/winnow-general-model
by InstaDeepAIWinnow recalibrates confidence scores and provides FDR control for de novo peptide sequencing (DNS) workflows. This repository hosts a pretrained, general-purpose calibrator that maps raw InstaNovo model confidences and complementary features (mass error, retention time, beam features, fragment…
InstaDeepAI/winnow-helaqc-model
by InstaDeepAIWinnow recalibrates confidence scores and provides FDR control for de novo peptide sequencing (DNS) workflows. This repository hosts a calibrator trained on the HeLa Single Shot dataset as referenced in our paper: De novo peptide sequencing rescoring and FDR estimation with Winnow.
shikunpunk/ask-dao
by shikunpunk> ⚠️ 重要:本仓库的 adapter 历史上因 PeftModel.frompretrained 双重包装导致 key 嵌套错误。 > 旧版本里 PeftModel.frompretrained 加载会"Found missing adapter keys"并静默丢弃全部权重, > 模型实际退化为 base Qwen2.5-3B-Instruct。 > 现在本仓库的 adapter_model.safetensors 已重新打包为标准深度 8(504/504 keys 命中),可被正确加载。 > 验证方式:见 shikunpunk/ask-dao-v0.3 仓库里 "Holdout…
antonknee/anisolv
by antonkneeshikunpunk/ask-dao-v0.2-2ep
by shikunpunk知识发现机器 —— 从生物医学论文推断「作者没有明说」的开放科学问题
Aurigene-AI/ChemFM-1B
by Aurigene-AISee the upstream model card for full details, training data and citation.
learning-unit/L1-30B-A5B
by learning-unitL1-30B-A5B is the Korean-locale medical foundation model from Lunit and Lunit Consortium. It is the 30B member of the L1 family, post-trained directly from Gravity-30B-A5B-Base, a sparse Mixture-of-Experts model developed by Trillion Labs and the Lunit Consortium.
Prior-Labs/tabpfn_3_5
by Prior-Labs### Model Overview TabPFN-3.5 is a transformer-based foundation model that uses in-context learning to solve tabular prediction problems in a forward pass. One checkpoint serves both classification and regression. Inference code can be found at https://github.com/PriorLabs/TabPFN.
UniParser/MolParser-Mobile
by UniParser💻 Github | 📄 Report | 🚀 Demo
UniParser/MolDetv2
by UniParserCompared to MolDet, our new MolDetv2 model leverages more manually annotated training data, with further optimizations specifically for reducing molecular false detections and improving bounding box regression, achieving stronger performance with a smaller model.
dinghhhhhhhhhhhhhhh/EvSpark
by dinghhhhhhhhhhhhhhhDrafter checkpoints for EvSpark: a small distilled drafter that accelerates single-stream Evo2 7B (StripedHyena2) generation while staying distributionally lossless.
ChongCong/Medical-SAM3
by ChongCongMedical SAM3 is a foundation model for universal prompt-driven medical image segmentation, obtained by fully fine-tuning SAM3 on large-scale, heterogeneous 2D and 3D medical imaging datasets with paired segmentation masks and text prompts.
egcastro/ReLSO
by egcastro# ReLSO: A Transformer-based Model for Latent Space Optimization and Generation of Proteins - github repo
Sentinal4D/PhenoSeq
by Sentinal4DPhenoSeq is a Gaussian diffusion model that generates scGPT RNA-seq embeddings conditioned on ViT-L microscopy imaging features. Given fluorescence microscopy images of a cell or well, it predicts a 512-dimensional scGPT embedding representing the transcriptomic state of individual cells — enabling…
vkali08/senba
by vkali08Senba is a source-anchored adaptation of MarS-FM (Valence Labs, ICLR 2026) for mdCATH backbone transitions at 450 K and one 50-frame lag. It is published as a validation-stage candidate together with the selection receipt, every comparison, and the preregistered protocols, so the claim below can be…
datamonkey/hyphaeon
by datamonkeyUltra-fast neural inference of episodic positive selection in molecular sequences.
tomatoai2026/TomatoPGFM
by tomatoai2026> TomatoPGFM v0.1.0 release metadata. Verify the SHA-256 value after downloading > the inference artifact before loading it.
nabbo/protenix_v2
by nabboThis repository provides the pretrained Protenix v2 model weights for protein structure prediction.
Paper: Arxiv | Website: Biomedica | Training instructions: OpenCLIP | Tutorial: Google Colab
SII-GAIR-NLP/RIBOSPAN-FM
by SII-GAIR-NLPMany full-length RNAs, particularly mRNAs, exceed the ~1K context lengths used to pretrain representative dense RNA encoders, forcing long transcripts to be truncated and preventing their 5′ UTR, CDS, and 3′ UTR from being modeled jointly at single-nucleotide resolution.
jlu-wsj/scRep
by jlu-wsjscRep is a PyTorch model for extracting cell embeddings from single-cell RNA-seq AnnData (.h5ad) inputs. This repository is a standalone Hugging Face release bundle: it includes checkpoint weights, the exact paired gene vocabulary, the inference implementation, and runnable examples.
mradermacher/Nidum-Gemma-2B-Uncensored-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
InstaDeepAI/instanovo-phospho-v1.0.0
by InstaDeepAIInstaNovo-P is a specialized transformer-based model for de novo peptide sequencing from phosphoproteomics mass spectrometry data. This model is specifically trained and optimized for identifying phosphorylated peptides and their modification sites.
datasetter458/bulk-chem-v1.0
by datasetter458A Data-Driven Router architecture for ADMET prediction.
NisargRhino/SmileBERTa
by NisargRhinoSmileBERTa is a RoBERTa-based language model for chemistry, built on the ChemBERTa architecture and fine-tuned to predict small-molecule fragment SMILES from full small-molecule drug SMILES.
Manhph2211/ECG-Scan
by Manhph2211litert-community/MedGemma-1.5-4B-IT
by litert-community![Figure #](
birder-project/rope_vit_reg8_b14_nps_avg_capi-dino-bio
by birder-projectA RoPE ViT Reg8 B/14 image encoder with average pooling, pretrained using CAPI-DINO on natural biological images. This model has not been fine-tuned for a specific classification task and is intended to be used as a general-purpose feature extractor or a backbone for downstream tasks like object…
Taykhoom/ModernGENA-large
by TaykhoomMinimal HuggingFace repackage of the large variant of ModernGENA -- a ModernBERT DNA encoder pretrained on vertebrate genomes with masked language modeling.
sciai-lab/boa
by sciai-labTrained checkpoints for the ICLR 2026 paper A Function-Centric Graph Neural Network Approach For Predicting Electron Densities.