polymathic-ai/MIMIC
https://huggingface.co/polymathic-ai/MIMICMIMIC is a multimodal encoder–decoder foundation model of the central dogma, trained jointly over DNA, RNA, and protein together with a range of structural and functional tracks. A single model embeds any subset of modalities into a shared representation space and generates any modality conditioned…
Sourced from
- HuggingFace — polymathic-ai/MIMIC
Related resources
Aquiles-ai/Evo2-1B-Base
by Aquiles-aiEvo2-1B-Base (Transformers port)
SII-GAIR-NLP/RIBOSPAN-10K-15
by SII-GAIR-NLPzhangtaolab/plant-dnagpt-BPE
by zhangtaolabThe plant DNA large language models (LLMs) contain a series of foundation models based on different model architectures, which are pre-trained on various plant reference genomes. All the models have a comparable model size between 90 MB and 150 MB, BPE tokenizer is used for tokenization and 8000…
This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…
SII-GAIR-NLP/RIBOSPAN-FM
by SII-GAIR-NLPMany full-length RNAs, particularly mRNAs, exceed the ~1K context lengths used to pretrain representative dense RNA encoders, forcing long transcripts to be truncated and preventing their 5′ UTR, CDS, and 3′ UTR from being modeled jointly at single-nucleotide resolution.
zhangtaolab/plant-dnabert-6mer
by zhangtaolabThe plant DNA large language models (LLMs) contain a series of foundation models based on different model architectures, which are pre-trained on various plant reference genomes. All the models have a comparable model size between 90 MB and 150 MB, BPE tokenizer is used for tokenization and 8000…