HantaBERT/HantaBERT

https://huggingface.co/HantaBERT/HantaBERT
Activeby HantaBERT01updated 3 months ago
Python

HantaBERT fine-tunes DNABERT-2 on hantavirus RNA sequences for three simultaneous classification tasks: species/lineage, host, and geographic origin. A single forward pass produces predictions for all three tasks along with a 768-dimensional embedding suitable for phylogenetic visualization.

Sourced from

  • HuggingFace — HantaBERT/HantaBERT

Related resources

Minimal HuggingFace repackage of the large variant of ModernGENA -- a ModernBERT DNA encoder pretrained on vertebrate genomes with masked language modeling.

Active541 month ago
Python
Active51 week ago
Python

This 1,120,772,224-parameter nucleotide-level causal language model is a member of the eight-model MarinDNA v0.5 parameter-scaling ladder developed with Marin. This repository contains only the final step-215573 checkpoint from run dna-bolinas-scaling-v0.5-h1920-p1B-0dc6f4, with its tokenizer…

Active2132 months ago
Python

The plant DNA large language models (LLMs) contain a series of foundation models based on different model architectures, which are pre-trained on various plant reference genomes. All the models have a comparable model size between 90 MB and 150 MB, BPE tokenizer is used for tokenization and 8000…

Idle3171 year ago
Python

Evo2-1B-Base (Transformers port)

Active1.7K3 weeks ago
Python

The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human…

Idle2.9K1 year ago
Python