Back

InterPET: A Curated Benchmark of Sequence Embeddings and Graph Architectures with Interpretability and Biological Validation for PETase Activity Prediction

Handrian, C.; Prakoso, I.

2026-08-23 bioinformatics
10.64898/2026.08.18.745360 bioRxiv
Show abstract

Motivation: Machine learning has emerged as a powerful accelerator for identifying PET-hydrolyzing enzymes (PETases). Yet, published models are often evaluated on benchmark performance alone, leaving their biological validity unexamined. Here we present InterPET, a curated benchmark and ablation study addressing both issues. Results: We aggregated sequences from four datasets (PlasticDB, PAZy, PlasticEnz, PEZY-miner), removing duplicate sequences, and filter data leakage, yielding a training set of 937 sequences and a benchmark of 139 sequences. Eight model configurations were trained and evaluated, spanning three embeddings (ESM-2, ProtT5, classical AAC/CTD descriptors), two tree-based classifiers (XGBoost, Random Forest), and two GraphSAGE variants differing in sequence-only and sequene plus 3D structure data. ESM-2 + XGBoost achieved the best performance (F1 = 0.91, AUC = 0.99, MCC = 0.90). SHAP-based feature attribution linked top-ranked AAC/CTD features (proline content, solvent accessibility, hydrophobicity) to known determinants of PETase activity, and cross-representation correlation showed that embedding-based models implicitly re-encode much of the same biophysical signal. However, in-silico mutagenesis revealed that the top-ranked M1 recovered only 0.5/3 catalytic-triad residues. These findings demonstrate that representation choice, classifier architecture, and evaluation criteria interact in ways a single leaderboard metric cannot capture. Availability and implementation: InterPET datasets and code are available at https://github.com/indiraprakoso/interpet/.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.