Find open-source science resources

A directory of tools, AI models, datasets, and research resources for biotech, bioinformatics, and other scientific fields. Aggregated from curated GitHub awesome-lists, HuggingFace, bio.tools, Bioconductor, and more.

881 of 7,064 resources

Showing 151–200

Nesso-1 is a fast, structure-based protein–ligand binding-affinity model. Given a protein sequence and a ligand (SMILES / CCD code / SDF), it predicts a binding affinity scalar along with a binder/non-binder score.

Active86.3K2 months ago

This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…

Active2132 months ago
Python

MarinDNA m5.1 is a 1.12B-parameter, nucleotide-level causal language model developed with Marin. This is the final m5.1 base-model checkpoint at step 59,158 from run dna-bolinas-mix-v0.9-p1B-i24-exp135-zoonomia-m5.1-bef41e, released with the A 1B standard Transformer rivals Evo 2 40B on variant…

Active1.7K2 months ago
Python

For a convenient overview and download list, visit our model page for this model.

Active5782 months ago
Python

Longevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. This checkpoint, L-Qwen3-0.6B, is the smallest family member and was produced by full-parameter supervised fine-tuning of Qwen/Qwen3-0.6B on aging-related multi-omics and…

Active902 months ago
Python

💻 GitHub Repo • 🤗 Model Collections • 📖 MemSFT Paper

Active1322 months ago

GenomeOcean-4B-v1.2 is a 4-billion-parameter causal language model for microbial genomic sequences. It is the June 2026 public release of the 4B v1.2 continued-training run, starting from GenomeOcean-4B and trained on an expanded corpus that adds IMGVR5 UViG, GTDB r226 representative genomes,…

Active1612 months ago

GenomeOcean-500M-v1.2 is a 500-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-500M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…

Active3112 months ago

GenomeOcean-100M-v1.2 is a 100-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-100M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…

Active1572 months ago

!Format !Task !Params !Type !License

Active5.9K2 months ago
Python

!Format !Task !Params !Type !License

Active4.6K2 months ago

!Format !Task !Params !Type !License

Active1162 months ago

HealthGPT-LoRA is a biomedical question-answering model built by fine-tuning Meta Llama 3.2 3B Instruct using QLoRA (PEFT) on the PubMedQA dataset.

Active222 months ago
Python

ProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.

Active192 months ago
Python

A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: smooth-L1 regression onto the frozen target, no negatives (the JEPA move).

Active02 months ago
Python

A DINOv2 ViT-S/14-reg fine-tuned so that an image of a molecular structure diagram embeds where its molecule embeds in the frozen MIST-28M embedding space. Objective: SigLIP sigmoid pairwise loss.

Active02 months ago
Python

A multilingual PII extractor for teams that need structured JSON from clinical and administrative text.

Active1382 months ago
Python

Contrastively fine-tuned ESM-2 150M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.

Active222 months ago
Python

🩺 HealthGPT-Pro: A High-Performance Multimodal Large Language Model for Medical Understanding and Analysis

Active2.1K2 months ago

A binary healthy hard coral vs bleached hard coral classifier built on top of the ReefNet species LoRA model BobDerBaum/bioclip-2.5-vith14-reefnet-lora, which provides the fine-tuned vision-encoder LoRA adapters. Only a small linear head is trained on top (frozen backbone + LoRA + 2-way linear…

Active02 months ago

A LoRA fine-tune of imageomics/bioclip-2.5-vith14 trained contrastively on the ReefNet 1.0 coral-reef species dataset (ReefNet/ReefNet-1.0), 92-class global curated split (train 48,312 / image-val 32,792 / image-test 33,090 / source-val 8,074; split cache 56ea94e36f9f).

Active42 months ago
Active02 months ago
Python

A lightweight language model designed to answer science questions clearly and accurately in English.

Active2202 months ago

!Nucleus Resynthesis

Active342 months ago

A foundation model over paired proteomic and transcriptomic expression. Attention between measured features yields a protein-protein association network without task-specific training.

Active02 months ago

For a convenient overview and download list, visit our model page for this model.

Active5792 months ago
Python
Active252 months ago

464 fine-tuned DNABERT models for regulatory variant effect prediction

Active242 months ago
Python
Active02 months ago
Python

GB.StructureEncoder is the encoder-only component of GB.StructureTokenizer for tokenization of protein structures.

Active742 months ago

GB.DNA-7B is DNA foundation model trained on 10.6 billion nucleotides from 796 species, enabling genome mining, in silico mutagenesis studies, gene expression prediction, and directed sequence generation.

Active872 months ago

polymathic-ai/MIMIC

by polymathic-ai

MIMIC is a multimodal encoder–decoder foundation model of the central dogma, trained jointly over DNA, RNA, and protein together with a range of structural and functional tracks. A single model embeds any subset of modalities into a shared representation space and generates any modality conditioned…

Active502 months ago

Trinity-Mini-AI-Scientist

Active132 months ago
Python

This is the first release of the Cherimoya Accessibility aTlas (CATv1): a collection of over 7,500 Cherimoya models trained on DNase-seq and ATAC-seq experiments from the ENCODE Project. Cherimoya models are state-of-the-art predictors of local chromatin accessibility, mapping a DNA sequence to a…

Active02 months ago

In search engines, rerankers are crucial for improving the accuracy of your retrieval system.

Active696.9K2 months ago
Python

In search engines, rerankers are crucial for improving the accuracy of your retrieval system.

Active1.3K2 months ago
Python

In retrieval systems, embedding models determine the quality of your search.

Active6.8K2 months ago
Python

OmniChem is a new series of large language models specialized for the domain of chemistry. It is designed to address the critical challenge of model hallucination in scientific applications. For OmniChem, we release this 7B instruction-tuned model with strong reasoning capabilities.

Active542 months ago

MoLFormer is a class of models pretrained on SMILES string representations of up to 1.1B molecules from ZINC and PubChem. This repository is for the model pretrained on 10% of both datasets.

Active255.3K2 months ago
Python

The commands below use the glyph package. Install it from the code repository:

Active02 months ago

The commands below use the glyph package. Install it from the code repository:

Active1932 months ago
Python

scUNVEIL is a pretrained foundation model for human single-cell RNA sequencing (scRNA-seq).

Active1482 months ago

This is a QLoRA adapter for query-focused structured extraction from one PubMed title and abstract. It was trained as part of BioEvidence Copilot and targets the repository's versioned ModelEvidenceExtraction JSON Schema.

Active142 months ago
Python

Healthcare Brain Procedure Surgery NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of surgical procedures, diagnostic tests, interventions, and procedural details from unstructured clinical text.

Active242 months ago
Python

Healthcare Brain Vitals NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of vital signs, body measurements, and physiological parameters from clinical text.

Active242 months ago
Python

Healthcare Brain Laboratory NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of laboratory test results, values, units, reference ranges, and abnormality flags from unstructured clinical text.

Active272 months ago
Python

Healthcare Brain Diagnosis ICD NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of diagnoses, conditions, and support for ICD-10/SNOMED code mapping from unstructured clinical text.

Active382 months ago
Python

Healthcare Brain Medication NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of medication names, dosages, routes, frequencies, and administration details from unstructured clinical text.

Active542 months ago
Python

Healthcare Brain Clinical Findings NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated extraction of clinical findings, diseases, conditions, anatomical locations, and clinical modifiers from unstructured clinical text.

Active02 months ago
Python

Healthcare Brain NER is a transformer-based clinical Named Entity Recognition model developed by Genzeon Platforms for automated detection and de-identification of Protected Health Information (PHI) and Personally Identifiable Information (PII) in unstructured clinical text.

Active312 months ago
Python