arcinstitute/evo2_20b

https://huggingface.co/arcinstitute/evo2_20b
Idleby arcinstitute10611updated 6 months ago

Evo 2 is a state-of-the-art DNA language model trained autoregressively on trillions of DNA tokens.

Sourced from

  • HuggingFacearcinstitute/evo2_20b

Related resources

Foundation models for genomics and transcriptomics pretrained on 3,000+ human genomes and 850+ diverse species, enabling chromatin accessibility prediction, splice site detection, and promoter classification across multiple model scales (InstaDeep, NVIDIA & TUM, Nature Methods 2023)

Idle9016 months ago
Jupyter Notebook
NOASSERTION

This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…

Active2131 month ago
Python

A PyTorch port of AlphaGenome, the DNA sequence model from Google DeepMind that predicts hundreds of genomic tracks at single base-pair resolution from sequences up to 1M bp.

Idle706 months ago

This repository contains the model presented in D3LM: A Discrete DNA Diffusion Language Model for Bidirectional DNA Understanding and Generation.

Idle326 months ago
Python

HantaBERT fine-tunes DNABERT-2 on hantavirus RNA sequences for three simultaneous classification tasks: species/lineage, host, and geographic origin. A single forward pass produces predictions for all three tasks along with a 768-dimensional embedding suitable for phylogenetic visualization.

Active02 months ago
Python

Pretrained LDARNet (~110M params) with learnable DNA tokenization (dynamic chunking + BiMamba-2).

Active02 months ago