songlab/tokenizer-dna-mlm
https://huggingface.co/songlab/tokenizer-dna-mlmSourced from
- HuggingFace — songlab/tokenizer-dna-mlm
Related resources
Foundation models for genomics and transcriptomics pretrained on 3,000+ human genomes and 850+ diverse species, enabling chromatin accessibility prediction, splice site detection, and promoter classification across multiple model scales (InstaDeep, NVIDIA & TUM, Nature Methods 2023)
duttaprat/DeepVRegulome
by duttaprat462 fine-tuned DNABERT models for regulatory variant effect prediction
Hengchang-Liu/D3LM-from-nt
by Hengchang-LiuThis repository contains the model presented in D3LM: A Discrete DNA Diffusion Language Model for Bidirectional DNA Understanding and Generation.
The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…
A PyTorch port of AlphaGenome, the DNA sequence model from Google DeepMind that predicts hundreds of genomic tracks at single base-pair resolution from sequences up to 1M bp.
DOEJGI/GenomeOcean-4B
by DOEJGIThis is the base model of GenomeOcean-4B. It is trained with Causal Language Modeling (CLM) and uses a BPE tokenizer with 4096 tokens. It supports a maximum sequence length of 10240 tokens (~50kbp).