DOEJGI/vhamster-models
https://huggingface.co/DOEJGI/vhamster-modelsSourced from
- HuggingFace — DOEJGI/vhamster-models
Related resources
DOEJGI/GenomeOcean-4B
by DOEJGIThis is the base model of GenomeOcean-4B. It is trained with Causal Language Modeling (CLM) and uses a BPE tokenizer with 4096 tokens. It supports a maximum sequence length of 10240 tokens (~50kbp).
DOEJGI/GenomeOcean-100M-v1.2
by DOEJGIGenomeOcean-100M-v1.2 is a 100-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-100M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…
DOEJGI/GenomeOcean-500M-v1.2
by DOEJGIGenomeOcean-500M-v1.2 is a 500-million-parameter causal language model for microbial genomic sequences. It is a continued-training checkpoint of GenomeOcean-500M (v1.0) trained on an expanded dataset that adds GTDB r226 representative genomes, INPHARED phage genomes, and the Zenodo RNA virus…
DOEJGI/GenomeOcean-4B-v1.2
by DOEJGIGenomeOcean-4B-v1.2 is a 4-billion-parameter causal language model for microbial genomic sequences. It is the June 2026 public release of the 4B v1.2 continued-training run, starting from GenomeOcean-4B and trained on an expanded corpus that adds IMGVR5 UViG, GTDB r226 representative genomes,…
This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…
vandijklab/C2S-Scale-Gemma-2-27B
by vandijklabGitHub homepage: Cell2Sentence GitHub