GrimSqueaker/ProtSent-V2-150M
https://huggingface.co/GrimSqueaker/ProtSent-V2-150MContrastively fine-tuned ESM-2 150M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
Sourced from
- HuggingFace — GrimSqueaker/ProtSent-V2-150M
Related resources
GrimSqueaker/ProtSent-V2-ESMC-300M
by GrimSqueakerContrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
GrimSqueaker/ProtSent-V2.5-35M
by GrimSqueakerProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.
This model card provides an overview of the intended use of the ESMC SAE models and examples of how to access them, but it does not have a specific model or model weights. To access each SAE model collection, use the links below:
biohub/ESMC-600M
by biohubESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…