GrimSqueaker/ProtSent-V2-ESMC-300M
https://huggingface.co/GrimSqueaker/ProtSent-V2-ESMC-300MContrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
Sourced from
- HuggingFace — GrimSqueaker/ProtSent-V2-ESMC-300M
Related resources
GrimSqueaker/ProtSent-V2-150M
by GrimSqueakerContrastively fine-tuned ESM-2 150M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
GrimSqueaker/ProtSent-V2.5-35M
by GrimSqueakerProtSent-V2 35M plus one more contrastive pass on a fresh draw of the corpus, with a DMS/ProteinGym CoSENT target and a Global Orthogonal Regularization term added.
FremyCompany/BioLORD-2023-M
by FremyCompany# FremyCompany/BioLORD-2023-M This model was trained using BioLORD, a new pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
biohub/ESMC-600M
by biohubESMC is a state-of-the-art protein language model that has learned the rules of protein biology from training on billions of protein sequences. ESMC provides representations of proteins enabling novel AI applications from therapeutic protein engineering to unlocking basic insights into protein…