How Not to be Seen: Predicting Unseen Enzyme Functions using Contrastive Learning
Ma, X.; Joshi, P.; Friedberg, I.; Li, Q.
Show abstract
MotivationPredicting enzyme function from its sequence is still an unsolved problem in the life sciences. Moreover, with the explosion of annotated genome data, we are inundated with potential enzymatic sequences that have not yet been biochemically characterized. While it is not possible to assign a not-yet-existing label to such a sequence, there is high value in placing the sequence as accurately as possible in known function space. Doing so can help provide more accurate falsifiable hypotheses for experimentalists wishing to characterize enzymes from specific functional families. ResultsHere we present a contrastive learning algorithm for predicting enzyme function from sequence. Our method, EnzPlacer, predicts the third, second, and first EC numbers for a protein whose fourth EC number is not in the training corpus. This novel prediction mechanism accurately places a protein sequence within a narrowed-down functional context, even if the precise function remains unknown. AvailabilityEnzPlacer is available from https://github.com/drxiangma/EnzPlacer under a GPL3 license. Contactqli@iastate.edu
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model 96%
- An adversarial scheme for integrating multi-modal data on protein function 96%
- DynamicGT: a dynamic-aware geometric transformer model to predict protein binding interfaces in flexible and disordered regions 95%
Similar papers in this journal
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 95%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 94%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.