DOEJGI/GenomeOcean-500M-v1.2

https://huggingface.co/DOEJGI/GenomeOcean-500M-v1.2
Activeby DOEJGI5251updated 1 month ago

GenomeOcean-500M-v1.2 is a 500-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-500M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…

Sourced from

  • HuggingFaceDOEJGI/GenomeOcean-500M-v1.2

Related resources

GenomeOcean-4B-v1.2 is a 4-billion-parameter causal language model for microbial genomic sequences. It is the June 2026 public release of the 4B v1.2 continued-training run, starting from GenomeOcean-4B and trained on an expanded corpus that adds IMGVR5 UViG, GTDB r226 representative genomes,…

Active1971 month ago

GenomeOcean-100M-v1.2 is a 100-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-100M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…

Active2061 month ago

This is the base model of GenomeOcean-4B. It is trained with Causal Language Modeling (CLM) and uses a BPE tokenizer with 4096 tokens. It supports a maximum sequence length of 10240 tokens (~50kbp).

Idle8581 year ago

Evo 2 is a state of the art DNA language model for long context modeling and design. Evo 2 models DNA sequences at single-nucleotide resolution at up to 1 million base pair context length using the StripedHyena 2 architecture, using Savanna.

Idle01 year ago

polymathic-ai/MIMIC

by polymathic-ai

MIMIC is a multimodal encoder–decoder foundation model of the central dogma, trained jointly over DNA, RNA, and protein together with a range of structural and functional tracks. A single model embeds any subset of modalities into a shared representation space and generates any modality conditioned…

Active501 month ago

This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…

Active2131 month ago
Python