Find open-source science resources
A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.
Filters
Health
Domain
Language
License
Source
Type(1)
881 of 7,064 resources
Showing 151–200
recursionpharma/nesso
by recursionpharmaNesso-1 is a fast, structure-based protein–ligand binding-affinity model. Given a protein sequence and a ligand (SMILES / CCD code / SDF), it predicts a binding affinity scalar along with a binder/non-binder score.
This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…
MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…
mradermacher/Gemma-2B-Uncensored-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
insilicomedicine/Qwen3-0.6B-Longevity
by insilicomedicineLongevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-0.6B, is the smallest family member and was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-0.6B on aging-related multi-omics and…
internlm/Intern-MemDec-4B
by internlm💻 GitHub Repo • 🤗 Model Collections • 📖 MemSFT Paper
DOEJGI/GenomeOcean-4B-v1.2
by DOEJGIGenomeOcean-4B-v1.2 is a 4-billion-parameter causal language model for microbial genomic sequences. It is the June 2026 public release of the 4B v1.2 continued-training run, starting from GenomeOcean-4B and trained on an expanded corpus that adds IMGVR5 UViG, GTDB r226 representative genomes,…
DOEJGI/GenomeOcean-500M-v1.2
by DOEJGIGenomeOcean-500M-v1.2 is a 500-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-500M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…
DOEJGI/GenomeOcean-100M-v1.2
by DOEJGIGenomeOcean-100M-v1.2 is a 100-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-100M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…
!Format !Task !Params !Type !License
!Format !Task !Params !Type !License
lemuralabs/Gemma-2B-Uncensored
by lemuralabs!Format !Task !Params !Type !License
llmithull/HealthGPT-LoRA
by llmithullHealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.
GrimSqueaker/ProtSent-V2.5-35M
by GrimSqueakerProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.
A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: smooth-L1 regression onto the frozen target, no negatives (the JEPA move).
A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: SigLIP sigmoid pairwise loss.
Meddies/meddies-pii
by MeddiesA multilingual PII extractor for teams that need structured JSON from clinical and administrative text.
GrimSqueaker/ProtSent-V2-150M
by GrimSqueakerContrastively fine-tuned ESM-2 150M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
🩺 HealthGPT-Pro: A High-Performance Multimodal Large Language Model for Medical Understanding and Analysis
A binary healthy hard coral vs bleached hard coral classifier built on top of the ReefNet species LoRA model BobDerBaum/bioclip-2.5-vith14-reefnet-lora, which provides the fine-tuned vision-encoder LoRA adapters. Only a small linear head is trained on top (frozen backbone + LoRA + 2-way linear…
A LoRA fine-tune of imageomics/bioclip-2.5-vith14 trained contrastively on the ReefNet 1.0 coral-reef species dataset (ReefNet/ReefNet-1.0), 92-class global curated split (train 48,312 / image-val 32,792 / image-test 33,090 / source-val 8,074; split cache 56ea94e36f9f).
ZeroOneAI/ZEO-Med-2
by ZeroOneAIHaamipromax/HamAI-Science-1b
by HaamipromaxA lightweight language model designed to answer science questions clearly and accurately in English.
rednaSander/omicsfm
by rednaSanderA foundation model over paired proteomic and transcriptomic expression. Attention between measured features yields a protein-protein association network without task-specific training.
mradermacher/BrainMed-8B-GGUF
by mradermacherFor a convenient overview and download list, visit our model page for this model.
duttaprat/DeepVRegulome
by duttaprat464 fine-tuned DNABERT models for regulatory variant effect prediction
PatSnap/Hiro-OCSR
by PatSnapgenbio-ai/GB.StructureEncoder
by genbio-aiGB.StructureEncoder is the encoder-only component of GB.StructureTokenizer for tokenization of protein structures.
genbio-ai/GB.DNA-7B
by genbio-aiGB.DNA-7B is DNA foundation model trained on 10.6 billion nucleotides from 796 species, enabling genome mining, in silico mutagenesis studies, gene expression prediction, and directed sequence generation.
polymathic-ai/MIMIC
by polymathic-aiMIMIC is a multimodal encoder–decoder foundation model of the central dogma, trained jointly over DNA, RNA, and protein together with a range of structural and functional tracks. A single model embeds any subset of modalities into a shared representation space and generates any modality conditioned…
Trinity-Mini-AI-Scientist
programmable-genomics/CATv1
by programmable-genomicsThis is the first release of the Cherimoya Accessibility aTlas (CATv1): a collection of over 7,500 Cherimoya models trained on DNase-seq and ATAC-seq experiments from the ENCODE Project. Cherimoya models are state-of-the-art predictors of local chromatin accessibility, mapping a DNA sequence to a…
zeroentropy/zerank-2-reranker
by zeroentropyIn search engines, rerankers are crucial for improving the accuracy of your retrieval system.
zeroentropy/zerank-1-reranker
by zeroentropyIn search engines, rerankers are crucial for improving the accuracy of your retrieval system.
zeroentropy/zembed-1-embedding
by zeroentropyIn retrieval systems, embedding models determine the quality of your search.
Billy-Liu-DUT/OmniChem-7B-v1
by Billy-Liu-DUTOmniChem is a new series of large language models specialized for the domain of chemistry. It is designed to address the critical challenge of model hallucination in scientific applications. For OmniChem, we release this 7B instruction-tuned model with strong reasoning capabilities.
ibm-research/MoLFormer-XL-both-10pct
by ibm-researchMoLFormer is a class of models pretrained on SMILES string representations of up to 1.1B molecules from ZINC and PubChem. This repository is for the model pretrained on 10% of both datasets.
EdisonScientific/OCSRGlyph
by EdisonScientificThe commands below use the glyph package. Install it from the code repository:
EdisonScientific/MarkushGlyph
by EdisonScientificThe commands below use the glyph package. Install it from the code repository:
thonzik/sc-unveil
by thonzikscUNVEIL is a pretrained foundation model for human single-cell RNA sequencing (scRNA-seq).
This is a QLoRA adapter for query-focused structured extraction from one PubMed title and abstract. It was trained as part of BioEvidence Copilot and targets the repository's versioned ModelEvidenceExtraction JSON Schema.
Healthcare Brain Procedure Surgery NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of surgical procedures, diagnostic tests, interventions, and procedural details from unstructured clinical text.
genzeonplatform/healthcare-brain-vitals-ner
by genzeonplatformHealthcare Brain Vitals NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of vital signs, body measurements, and physiological parameters from clinical text.
genzeonplatform/healthcare-brain-laboratory-ner
by genzeonplatformHealthcare Brain Laboratory NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of laboratory test results, values, units, reference ranges, and abnormality flags from unstructured clinical text.
Healthcare Brain Diagnosis ICD NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of diagnoses, conditions, and support for ICD-10/SNOMED code mapping from unstructured clinical text.
genzeonplatform/healthcare-brain-medication-ner
by genzeonplatformHealthcare Brain Medication NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of medication names, dosages, routes, frequencies, and administration details from unstructured clinical text.
Healthcare Brain Clinical Findings NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of clinical findings, diseases, conditions, anatomical locations, and clinical modifiers from unstructured clinical text.
genzeonplatform/healthcare-brain-ner
by genzeonplatformHealthcare Brain NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated detection and de-identification of Protected Health Information (PHI) and Personally Identifiable Information (PII) in unstructured clinical text.