DOEJGI/vhamster-models

https://huggingface.co/DOEJGI/vhamster-models
Activeby DOEJGI01updated 1 week ago

Sourced from

  • HuggingFace — DOEJGI/vhamster-models

Related resources

This is the base model of GenomeOcean-4B. It is trained with Causal Language Modeling (CLM) and uses a BPE tokenizer with 4096 tokens. It supports a maximum sequence length of 10240 tokens (~50kbp).

Idle8581 year ago

GenomeOcean-100M-v1.2 is a 100-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-100M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…

Active1881 month ago

GenomeOcean-500M-v1.2 is a 500-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-500M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…

Active5081 month ago

GenomeOcean-4B-v1.2 is a 4-billion-parameter causal language model for microbial genomic sequences. It is the June 2026 public release of the 4B v1.2 continued-training run, starting from GenomeOcean-4B and trained on an expanded corpus that adds IMGVR5 UViG, GTDB r226 representative genomes,…

Active2011 month ago

This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…

Active2131 month ago
Python

GitHub homepage: Cell2Sentence GitHub

Idle1.5K10 months ago
Python