aasatorres/esm2-sae-topk-16384-k512
https://huggingface.co/aasatorres/esm2-sae-topk-16384-k512Sparse Autoencoder (SAE) trained on residue-level embeddings from ESM-2 (650M, layer 33) for interpretability research on protein language models.
Sourced from
- HuggingFace — aasatorres/esm2-sae-topk-16384-k512
Related resources
This model may be overfit to some extent (see below). Try running this notebook on the datasets linked to in the notebook. See if you can figure out why the metrics differ so much on the datasets. Is it due to something like sequence similarity in the train/test split?
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
AmelieSchreiber/esm_interact
by AmelieSchreiberThis model was finetuned on concatenated pairs of interacting proteins in much the same way as PepMLM. It is meant to generate interaction partners for proteins using the masked language modeling capabilities of ESM-2. The model is not well tested, so use with caution.
RomeroLab-Duke/prism-antibody
by RomeroLab-DukePRISM is an antibody language model that jointly predicts amino acid identity and germline/non-germline (GL/NGL) position classification, enabling developability-aware antibody sequence modeling.
biohub/ESMC-6B
by biohubESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…
westlake-repl/Evolla-10B
by westlake-replA frontier protein-language generative model — because proteins deserve better small talk.