Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source
Type(1)
884 of 7,068 resources
Showing 101–150
SII-GAIR-NLP/RIBOSPAN-FM
by SII-GAIR-NLPMany full-length RNAs, particularly mRNAs, exceed the ~1K context lengths used to pretrain representative dense RNA encoders, forcing long transcripts to be truncated and preventing their 5′ UTR, CDS, and 3′ UTR from being modeled jointly at single-nucleotide resolution.
jlu-wsj/scRep
by jlu-wsjscRep is a PyTorch model for extracting cell embeddings from single-cell RNA-seq AnnData (.h5ad) inputs. This repository is a standalone Hugging Face release bundle: it includes checkpoint weights, the exact paired gene vocabulary, the inference implementation, and runnable examples.
mradermacher/Nidum-Gemma-2B-Uncensored-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
InstaDeepAI/instanovo-phospho-v1.0.0
by InstaDeepAIInstaNovo-P is a specialized transformer-based model for de novo peptide sequencing from phosphoproteomics mass spectrometry data. This model is specifically trained and optimized for identifying phosphorylated peptides and their modification sites.
datasetter458/bulk-chem-v1.0
by datasetter458A Data-Driven Router architecture for ADMET prediction.
NisargRhino/SmileBERTa
by NisargRhinoSmileBERTa is a RoBERTa-based language model for chemistry, built on the ChemBERTa architecture and fine-tuned to predict small-molecule fragment SMILES from full small-molecule drug SMILES.
Manhph2211/ECG-Scan
by Manhph2211litert-community/MedGemma-1.5-4B-IT
by litert-community.
SII-GAIR-NLP/RIBOSPAN-10K-15
by SII-GAIR-NLPpregH/MolecularDiffusion
by pregHCENO-1B-1m is a checkpoint of the CENO base DNA foundation model (1M context (stage 4)). It is a plain causal language model over genomic sequence on a Nemotron-H Mamba/Attention/MoE hybrid backbone, with no MSA inputs.
omtx/lula-2
by omtxomtx/lula-1.1
by omtxomtx/lula-1
by omtxJerrryNie/ConceptCLIP
by JerrryNieloveCloud/OmniTCR
by loveCloudOmniTCR is a component-aware autoregressive foundation model for learning relationships among peptide epitopes, major histocompatibility complex (MHC) molecules, T-cell receptor alpha chains (TRA) and T-cell receptor beta chains (TRB).
tzcfly/PertMind
by tzcflyPertMind is a biological language model built around a central discovery: public cellular perturbation atlases can be reorganized into reinforcement-learning environments, where measured gene responses act as computable reward signals for biological reasoning.
NavitraTechnologies01/navikinase-1.0
by NavitraTechnologies01A from-scratch, decoder-only protein language model for the phosphotransferase superfamily (EC 2.7.-: protein kinases plus sugar/lipid/nucleotide kinases), trained entirely locally on Apple Silicon via MLX — no cloud compute, no fine-tuning of an existing model.
Run locally - Benchmarks - Whitepaper - Model details - Responsible use
NOC-Lab/AbCDR-ESMC
by NOC-LabThis model is a fine-tuned version of ESMC-600M (ESM Cambrian) for paired antibody variable-domain sequences containing heavy and light chains. It was trained using a CDR-focused masking strategy to improve representations for antibody binding affinity prediction.
Native OpenMed and OpenMedKit vision-language inference for Apple Silicon, including local clinical-document and chart workflows on Mac, iPhone, and iPad.
Eleuthia is a binary protein variant classifier that predicts whether a mutated protein sequence is likely Pathogenic or Benign. It is fine-tuned from ESM-2 and intended for research support in variant prioritization workflows.
ratschlab/DeepSpotM
by ratschlabA 350M encoder that finds nine types of personally identifiable information across 17 languages and returns exact character spans for review and redaction.
IQuestLab/IQuest-UBio-MolFM-V1.5
by IQuestLabUBio-MolFM is a foundation-model suite for molecular modeling, designed for bio-systems. This release, UBio-MolFM-V1.5 (Stage 3), is built on the E2Former-V2 linear-scaling equivariant transformer and is the checkpoint used for every simulation reported in UBio-MolFM: Enabling Biomolecular Dynamics…
dmolino/text2ct-weights
by dmolinoCheckpoints for "From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation" (Molino et al., BMVC 2026).
prathmeshadsod/BondShift-Llama-3.3-70B-Instruct
by prathmeshadsodBondShift: Organic Mechanism Reasoning
Parrotlet-a 2.5 Pro is a purpose-built automatic speech recognition (ASR) model for medical speech in Indian healthcare settings. It transcribes Indian English, Hindi, Marathi, Kannada and Telugu, including the heavily code-mixed speech typical of real consultations (English drug names and clinical…
Vision-language model for dermatology, pretrained with MAGEN (Multi-Agent data GENeration) and O-MAKE (Ontology-based Multi-Aspect Knowledge-Enhanced pretraining).
Neurazum/VLbai-2.6AD
by NeurazumA clinical reasoning assistant for early-stage Alzheimer's assessment. It joins a 3D MRI + biomarker classifier (Vbai-2.6AD) to a reasoning LLM (Gemma 4 12B) inside a single forward pass — the diagnosis is passed as vectors, not text.
Trinetralab/matgraph-cli
by TrinetralabA modern CLI and GraphQL API tool for Material Science Deep Learning Pipelines.
GrimSqueaker/ProtSent-V2-ESMC-300M
by GrimSqueakerContrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
PypCoder/SERAPH
by PypCoderSERAPH is a deep learning model designed for 3-state (Q3) protein secondary structure prediction. It processes raw single amino acid sequences and predicts residue-level secondary structure states: Alpha Helix (H), Beta Sheet (E), or Coil/Loop (C).
!TVBP Architecture !Parameters !Trainable !Brownian Reservoir !Framework !Biology
prov-gigatime/gigatime-flash
by prov-gigatimeprov-gigatime/GigaTIME
by prov-gigatimeprov-gigapath/prov-gigapath-flash
by prov-gigapathprov-gigapath/prov-gigapath
by prov-gigapathAignostics/RudolfV-2-S
by AignosticsAignostics/RudolfV-2-B
by AignosticsAignostics/RudolfV-2
by AignosticsHuggingFaceBio/Carbon-3B
by HuggingFaceBioTechnical Report 🧬