aasatorres/esm2-sae-topk-16384-k512

https://huggingface.co/aasatorres/esm2-sae-topk-16384-k512
Activeby aasatorres182updated 2 months ago

Sparse Autoencoder (SAE) trained on residue-level embeddings from ESM-2 (650M, layer 33) for interpretability research on protein language models.

Sourced from

  • HuggingFaceaasatorres/esm2-sae-topk-16384-k512

Related resources

This model may be overfit to some extent (see below). Try running this notebook on the datasets linked to in the notebook. See if you can figure out why the metrics differ so much on the datasets. Is it due to something like sequence similarity in the train/test split?

Stale312 years ago
Python

This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:

Active01 month ago
Python

This model was finetuned on concatenated pairs of interacting proteins in much the same way as PepMLM. It is meant to generate interaction partners for proteins using the masked language modeling capabilities of ESM-2. The model is not well tested, so use with caution.

Stale82 years ago
Python

PRISM is an antibody language model that jointly predicts amino acid identity and germline/non-germline (GL/NGL) position classification, enabling developability-aware antibody sequence modeling.

Active01 month ago

ESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…

Active1.3M1 month ago
Python

A frontier protein-language generative model — because proteins deserve better small talk.

Active163 months ago