darlednik/LDARNet-2M

https://huggingface.co/darlednik/LDARNet-2M
Activeby darlednik01updated 3 months ago

Pretrained LDARNet (~2M params) with learnable DNA tokenization (dynamic chunking + BiMamba-2).

Sourced from

  • HuggingFace — darlednik/LDARNet-2M

Related resources

Pretrained LDARNet (~110M params) with learnable DNA tokenization (dynamic chunking + BiMamba-2).

Active03 months ago

# ModernGENA base ModernGENA is a DNA foundation model based on ModernBERT (a modernized BERT-style encoder architecture) adapted for genomic sequence modeling. ModernGENA base is the 377M-parameter version introduced in the paper Back to BERT in 2026: ModernGENA as a Strong, Efficient Baseline for…

Active4285 months ago

The plant DNA large language models (LLMs) contain a series of foundation models based on different model architectures, which are pre-trained on various plant reference genomes. All the models have a comparable model size between 90 MB and 150 MB, BPE tokenizer is used for tokenization and 8000…

Idle41 year ago

Minimal HuggingFace repackage of the large variant of ModernGENA -- a ModernBERT DNA encoder pretrained on vertebrate genomes with masked language modeling.

Active541 month ago
Python
Active03 months ago

The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…

Idle17.9K1 year ago
Python