CLEAN
Predicts enzyme function (EC numbers) from protein sequence
- Field
- Biology & genomics
- Method
- Deep learning
- Data it takes
- sequence
- License
- Research use only — restricted use
CLEAN assigns Enzyme Commission numbers to protein sequences using contrastive learning. It outperforms similarity-based annotation on enzymes with no well-characterized close relative, which describes most sequences recovered from metagenomes.
It runs on CPU and scales to large sequence sets.
The license permits research use only.
Catalog entry last checked 2026-08-17. All models →
Run CLEAN on your data
The lab runs this on its own compute, in a container, with the inputs and parameters recorded alongside the result. Initial scoping conversations are free.
Contact the Lab