Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

884 of 7,068 resources

Showing 101–150

Many full-length RNAs, particularly mRNAs, exceed the ~1K context lengths used to pretrain representative dense RNA encoders, forcing long transcripts to be truncated and preventing their 5′ UTR, CDS, and 3′ UTR from being modeled jointly at single-nucleotide resolution.

Active01 month ago

scRep is a PyTorch model for extracting cell embeddings from single-cell RNA-seq AnnData (.h5ad) inputs. This repository is a standalone Hugging Face release bundle: it includes checkpoint weights, the exact paired gene vocabulary, the inference implementation, and runnable examples.

Active01 month ago

For a convenient overview and download list, visit our model page for this model.

Active9741 month ago
Python

InstaNovo-P is a specialized transformer-based model for de novo peptide sequencing from phosphoproteomics mass spectrometry data. This model is specifically trained and optimized for identifying phosphorylated peptides and their modification sites.

Active1021 month ago

A Data-Driven Router architecture for ADMET prediction.

Active01 month ago

SmileBERTa is a RoBERTa-based language model for chemistry, built on the ChemBERTa architecture and fine-tuned to predict small-molecule fragment SMILES from full small-molecule drug SMILES.

Active561 month ago
Python
Active241 month ago
Python
Active3831 month ago

![Figure #](

Active10.4K1 month ago

A RoPE ViT Reg8 B/14 image encoder with average pooling, pretrained using CAPI-DINO on natural biological images. This model has not been fine-tuned for a specific classification task and is intended to be used as a general-purpose feature extractor or a backbone for downstream tasks like object…

Active281 month ago

Minimal HuggingFace repackage of the large variant of ModernGENA -- a ModernBERT DNA encoder pretrained on vertebrate genomes with masked language modeling.

Active541 month ago
Python

sciai-lab/boa

by sciai-lab

Trained checkpoints for the ICLR 2026 paper A Function-Centric Graph Neural Network Approach For Predicting Electron Densities.

Active01 month ago

This repository provides four SCIAI-DFT models for density optimization and geometry optimization with machine-learned orbital-free density functional theory (ML-OFDFT).

Active01 month ago
Active11 month ago

!Aerova

Active1201 month ago
Python
Active01 month ago

CENO-1B-1m is a checkpoint of the CENO base DNA foundation model (1M context (stage 4)). It is a plain causal language model over genomic sequence on a Nemotron-H Mamba/Attention/MoE hybrid backbone, with no MSA inputs.

Active4931 month ago
Python

!Aerova

Active251 month ago
Python
Active01 month ago
Active01 month ago
Active101 month ago
Active7611 month ago
Python

OmniTCR is a component-aware autoregressive foundation model for learning relationships among peptide epitopes, major histocompatibility complex (MHC) molecules, T-cell receptor alpha chains (TRA) and T-cell receptor beta chains (TRB).

Active01 month ago
Python

PertMind is a biological language model built around a central discovery: public cellular perturbation atlases can be reorganized into reinforcement-learning environments, where measured gene responses act as computable reward signals for biological reasoning.

Active5031 month ago
Python

A from-scratch, decoder-only protein language model for the phosphotransferase superfamily (EC 2.7.-: protein kinases plus sugar/lipid/nucleotide kinases), trained entirely locally on Apple Silicon via MLX — no cloud compute, no fine-tuning of an existing model.

Active821 month ago
Python

Run locally - Benchmarks - Whitepaper - Model details - Responsible use

Active2.1K1 month ago

This model is a fine-tuned version of ESMC-600M (ESM Cambrian) for paired antibody variable-domain sequences containing heavy and light chains. It was trained using a CDR-focused masking strategy to improve representations for antibody binding affinity prediction.

Active251 month ago

Native OpenMed and OpenMedKit vision-language inference for Apple Silicon, including local clinical-document and chart workflows on Mac, iPhone, and iPad.

Active2371 month ago

!OpenDDE banner

Active23K1 month ago

Eleuthia is a binary protein variant classifier that predicts whether a mutated protein sequence is likely Pathogenic or Benign. It is fine-tuned from ESM-2 and intended for research support in variant prioritization workflows.

Active621 month ago
Python
Active6961 month ago

A 350M encoder that finds nine types of personally identifiable information across 17 languages and returns exact character spans for review and redaction.

Active8341 month ago
Python

UBio-MolFM is a foundation-model suite for molecular modeling, designed for bio-systems. This release, UBio-MolFM-V1.5 (Stage 3), is built on the E2Former-V2 linear-scaling equivariant transformer and is the checkpoint used for every simulation reported in UBio-MolFM: Enabling Biomolecular Dynamics…

Active91 month ago

Checkpoints for "From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation" (Molino et al., BMVC 2026).

Active431 month ago

BondShift: Organic Mechanism Reasoning

Active131 month ago
Python

Parrotlet-a 2.5 Pro is a purpose-built automatic speech recognition (ASR) model for medical speech in Indian healthcare settings. It transcribes Indian English, Hindi, Marathi, Kannada and Telugu, including the heavily code-mixed speech typical of real consultations (English drug names and clinical…

Active1.2K2 months ago
Python

Vision-language model for dermatology, pretrained with MAGEN (Multi-Agent data GENeration) and O-MAKE (Ontology-based Multi-Aspect Knowledge-Enhanced pretraining).

Active592 months ago

A clinical reasoning assistant for early-stage Alzheimer's assessment. It joins a 3D MRI + biomarker classifier (Vbai-2.6AD) to a reasoning LLM (Gemma 4 12B) inside a single forward pass — the diagnosis is passed as vectors, not text.

Active02 months ago
Python

A modern CLI and GraphQL API tool for Material Science Deep Learning Pipelines.

Active02 months ago

Contrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.

Active452 months ago
Python

SERAPH is a deep learning model designed for 3-state (Q3) protein secondary structure prediction. It processes raw single amino acid sequences and predicts residue-level secondary structure states: Alpha Helix (H), Beta Sheet (E), or Coil/Loop (C).

Active02 months ago

!TVBP Architecture !Parameters !Trainable !Brownian Reservoir !Framework !Biology

Active1202 months ago
Active2772 months ago
Active1182 months ago
Active43.2K2 months ago
Python
Active1.3K2 months ago
Python

Technical Report 🧬

Active5.2K2 months ago
Python