ESM C

Creates protein-sequence embeddings for tasks such as function and property prediction.

Field
Biology & genomics
Method
Foundation model
Data it takes
sequence
License
MIT

ESM C is a protein language model producing per-residue and whole-sequence embeddings. These feed downstream predictors of function, stability and localization, and generally outperform hand-designed sequence features.

It is the default protein representation model in this catalog. The license is MIT and the weights are ungated.

Upstream project →

Catalog entry last checked 2026-08-17. All models →

Run ESM C on your data

The lab runs this on its own compute, in a container, with the inputs and parameters recorded alongside the result. Initial scoping conversations are free.

Contact the Lab