Homology-Based Variant-Effect Predictors Break Down on Cytochrome P450 Pharmacogenes
Xu, H.; Samori, I.; Nayar, G.; Altman, R. B.
Show abstract
Cytochrome P450 (CYP) enzymes metabolize roughly three-quarters of clinically used drugs; genetic variation in these enzymes is a leading source of interindividual differences in drug response. Predicting a variants functional effect is therefore critical, yet the consequences of most CYP variants remain unknown. Many state-of-the-art variant-effect predictors rest on a homology-based paradigm that scores variants by evolutionary conservation -- an assumption that pharmacogenes including CYPs violate. Indeed, focusing on human CYPs, we show that homology-based models fail systematically. AlphaMissense (AM) assigns variants to its "ambiguous" class at nearly twice the proteome-wide rate across six CYPs, and within that class the scores are essentially uncorrelated with CYP2C9 DMS activity (Spearmans{rho} = 0.069); Evolutionary Scale Modeling 2 (ESM-2) shows the same pattern. We further hypothesized that non-homology-based features (sequence position, substitution chemistry, binding-site distance, and secondary structure) might help resolve the ambiguous calls, but their explanatory power is weak. Using CYP2C9 DMS activity as ground truth, we built a k-nearest-neighbors model over ESM-2 embeddings and ensembled it with AM and ESM-2 masked marginal probability, improving the ambiguous-class correlation roughly ten-fold, from{rho} = 0.069 to 0.715 (overall{rho} from 0.638 to 0.825). However, this markedly improved accuracy does not translate into agreement with clinical annotations. Drawing on evidence that a variants effect can depend on the drug, we hypothesize that substrate identity is the key missing feature in current models, and that predicting function for these multi-substrate enzymes may require redefining function as substrate-conditioned.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Transfer learning enables prediction of CYP2D6 haplotype function 95%
- Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP) 92%
- Assessing variant effect predictors and disease mechanisms in intrinsically disordered proteins 91%
Similar papers in this journal
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 92%
- BindPred: A Framework for Predicting Protein-Protein Binding Affinity from Language Model Embeddings 92%
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.