Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

869 of 7,050 resources

Showing 51–100

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active22K3 weeks ago
Python

ESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…

Active208.7K3 weeks ago
Python

ESMFold2 is a state-of-the-art model for protein structure prediction and design that defines a new frontier for speed and accuracy. The model predicts high-resolution, all-atom 3D protein structures directly from amino acid sequences, with optional multiple sequence alignment (MSA) input for…

Active181.7K3 weeks ago
Python
Active43 weeks ago
Python

A domain-adapted small language model specialized for molecular conformation analysis, computational biochemistry, and pharmacology reasoning.

Active963 weeks ago
Python

Evo2-7B (Transformers port)

Active1.3K3 weeks ago
Python

Evo2-1B-Base (Transformers port)

Active1.7K3 weeks ago
Python

This bundle contains standardized CSV files and a Jupyter sample notebook for RMMol frozen-embedding examples across Biophysics, Physiological, Physical Chemistry, and Quantum Mechanics.

Active03 weeks ago

Machine-learned orbital-free density functional theory

Active03 weeks ago

CycleGAN generators that synthesise Ki-67 and pHH3 immunohistochemistry (IHC) appearance from H&E histopathology tiles of triple-negative breast cancer (TNBC).

Active03 weeks ago

!Benchmark card: Hertz 0.7F vs same-size models

Active603 weeks ago

# Geneformer Geneformer is a foundational transformer model pretrained on a large-scale corpus of human single cell transcriptomes to enable context-aware predictions in settings with limited data in network biology.

Active4.6K3 weeks ago
Python

ONNX conversions of the existing RxnScribe, MolScribe and English EasyOCR checkpoints. This is a community conversion, not a newly trained model or an upstream release. Use the matching RxnScribe feature branch and its RxnScribeONNX interface. PyTorch is needed for export, not inference.

Active03 weeks ago
Active03 weeks ago

- Project page - Try Packora (live demo) - Paper - Code - Dataset manifests - Hugging Face collection

Active03 weeks ago

Winnow recalibrates confidence scores and provides FDR control for de novo peptide sequencing (DNS) workflows. This repository hosts a pretrained, general-purpose calibrator that maps raw InstaNovo model confidences and complementary features (mass error, retention time, beam features, fragment…

Active553 weeks ago

Winnow recalibrates confidence scores and provides FDR control for de novo peptide sequencing (DNS) workflows. This repository hosts a calibrator trained on the HeLa Single Shot dataset as referenced in our paper: De novo peptide sequencing rescoring and FDR estimation with Winnow.

Active513 weeks ago

> ⚠️ 重要:本仓库的 adapter 历史上因 PeftModel.frompretrained 双重包装导致 key 嵌套错误。 > 旧版本里 PeftModel.frompretrained 加载会"Found missing adapter keys"并静默丢弃全部权重, > 模型实际退化为 base Qwen2.5-3B-Instruct。 > 现在本仓库的 adapter_model.safetensors 已重新打包为标准深度 8(504/504 keys 命中),可被正确加载。 > 验证方式:见 shikunpunk/ask-dao-v0.3 仓库里 "Holdout…

Active783 weeks ago
Python
Active03 weeks ago

知识发现机器 —— 从生物医学论文推断「作者没有明说」的开放科学问题

Active03 weeks ago
Python

See the upstream model card for full details, training data and citation.

Active2153 weeks ago
Python

L1-30B-A5B is the Korean-locale medical foundation model from Lunit and Lunit Consortium. It is the 30B member of the L1 family, post-trained directly from Gravity-30B-A5B-Base, a sparse Mixture-of-Experts model developed by Trillion Labs and the Lunit Consortium.

Active3873 weeks ago
Python

### Model Overview TabPFN-3.5 is a transformer-based foundation model that uses in-context learning to solve tabular prediction problems in a forward pass. One checkpoint serves both classification and regression. Inference code can be found at https://github.com/PriorLabs/TabPFN.

Active10.2K4 weeks ago

💻 Github | 📄 Report | 🚀 Demo

Active2664 weeks ago
Python

Compared to MolDet, our new MolDetv2 model leverages more manually annotated training data, with further optimizations specifically for reducing molecular false detections and improving bounding box regression, achieving stronger performance with a smaller model.

Active04 weeks ago

dinghhhhhhhhhhhhhhh/EvSpark

by dinghhhhhhhhhhhhhhh

Drafter checkpoints for EvSpark: a small distilled drafter that accelerates single-stream Evo2 7B (StripedHyena2) generation while staying distributionally lossless.

Active04 weeks ago

Medical SAM3 is a foundation model for universal prompt-driven medical image segmentation, obtained by fully fine-tuning SAM3 on large-scale, heterogeneous 2D and 3D medical imaging datasets with paired segmentation masks and text prompts.

Active44 weeks ago

# ReLSO: A Transformer-based Model for Latent Space Optimization and Generation of Proteins - github repo

Active01 month ago

PhenoSeq is a Gaussian diffusion model that generates scGPT RNA-seq embeddings conditioned on ViT-L microscopy imaging features. Given fluorescence microscopy images of a cell or well, it predicts a 512-dimensional scGPT embedding representing the transcriptomic state of individual cells — enabling…

Active601 month ago

Senba is a source-anchored adaptation of MarS-FM (Valence Labs, ICLR 2026) for mdCATH backbone transitions at 450 K and one 50-frame lag. It is published as a validation-stage candidate together with the selection receipt, every comparison, and the preregistered protocols, so the claim below can be…

Active01 month ago

Ultra-fast neural inference of episodic positive selection in molecular sequences.

Active3.1K1 month ago

> TomatoPGFM v0.1.0 release metadata. Verify the SHA-256 value after downloading > the inference artifact before loading it.

Active241 month ago

This repository provides the pretrained Protenix v2 model weights for protein structure prediction.

Active01 month ago

!SynthVision

Active13.9K1 month ago

Paper: Arxiv     |     Website: Biomedica     |     Training instructions: OpenCLIP     |     Tutorial: Google Colab

Active01 month ago

Many full-length RNAs, particularly mRNAs, exceed the ~1K context lengths used to pretrain representative dense RNA encoders, forcing long transcripts to be truncated and preventing their 5′ UTR, CDS, and 3′ UTR from being modeled jointly at single-nucleotide resolution.

Active01 month ago

scRep is a PyTorch model for extracting cell embeddings from single-cell RNA-seq AnnData (.h5ad) inputs. This repository is a standalone Hugging Face release bundle: it includes checkpoint weights, the exact paired gene vocabulary, the inference implementation, and runnable examples.

Active01 month ago

For a convenient overview and download list, visit our model page for this model.

Active8031 month ago
Python

InstaNovo-P is a specialized transformer-based model for de novo peptide sequencing from phosphoproteomics mass spectrometry data. This model is specifically trained and optimized for identifying phosphorylated peptides and their modification sites.

Active1231 month ago

A Data-Driven Router architecture for ADMET prediction.

Active01 month ago

SmileBERTa is a RoBERTa-based language model for chemistry, built on the ChemBERTa architecture and fine-tuned to predict small-molecule fragment SMILES from full small-molecule drug SMILES.

Active561 month ago
Python
Active241 month ago
Python
Active3831 month ago

![Figure #](

Active10.3K1 month ago

A RoPE ViT Reg8 B/14 image encoder with average pooling, pretrained using CAPI-DINO on natural biological images. This model has not been fine-tuned for a specific classification task and is intended to be used as a general-purpose feature extractor or a backbone for downstream tasks like object…

Active281 month ago

Minimal HuggingFace repackage of the large variant of ModernGENA -- a ModernBERT DNA encoder pretrained on vertebrate genomes with masked language modeling.

Active541 month ago
Python

sciai-lab/boa

by sciai-lab

Trained checkpoints for the ICLR 2026 paper A Function-Centric Graph Neural Network Approach For Predicting Electron Densities.

Active01 month ago