Back

How Not to be Seen: Predicting Unseen Enzyme Functions using Contrastive Learning

Ma, X.; Joshi, P.; Friedberg, I.; Li, Q.

2026-02-24 bioinformatics
10.64898/2026.02.23.707489 bioRxiv
Show abstract

MotivationPredicting enzyme function from its sequence is still an unsolved problem in the life sciences. Moreover, with the explosion of annotated genome data, we are inundated with potential enzymatic sequences that have not yet been biochemically characterized. While it is not possible to assign a not-yet-existing label to such a sequence, there is high value in placing the sequence as accurately as possible in known function space. Doing so can help provide more accurate falsifiable hypotheses for experimentalists wishing to characterize enzymes from specific functional families. ResultsHere we present a contrastive learning algorithm for predicting enzyme function from sequence. Our method, EnzPlacer, predicts the third, second, and first EC numbers for a protein whose fourth EC number is not in the training corpus. This novel prediction mechanism accurately places a protein sequence within a narrowed-down functional context, even if the precise function remains unknown. AvailabilityEnzPlacer is available from https://github.com/drxiangma/EnzPlacer under a GPL3 license. Contactqli@iastate.edu

Published in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.