CLEAN

Predicts enzyme function (EC numbers) from protein sequence

Field
Biology & genomics
Method
Deep learning
Data it takes
sequence
License
Research use only — restricted use

CLEAN assigns Enzyme Commission numbers to protein sequences using contrastive learning. It outperforms similarity-based annotation on enzymes with no well-characterized close relative, which describes most sequences recovered from metagenomes.

It runs on CPU and scales to large sequence sets.

The license permits research use only.

Upstream project →

Catalog entry last checked 2026-08-17. All models →

Run CLEAN on your data

The lab runs this on its own compute, in a container, with the inputs and parameters recorded alongside the result. Initial scoping conversations are free.

Contact the Lab