Back

Homology-Based Variant-Effect Predictors Break Down on Cytochrome P450 Pharmacogenes

Xu, H.; Samori, I.; Nayar, G.; Altman, R. B.

2026-08-05 bioinformatics
10.64898/2026.07.30.741616 bioRxiv
Show abstract

Cytochrome P450 (CYP) enzymes metabolize roughly three-quarters of clinically used drugs; genetic variation in these enzymes is a leading source of interindividual differences in drug response. Predicting a variants functional effect is therefore critical, yet the consequences of most CYP variants remain unknown. Many state-of-the-art variant-effect predictors rest on a homology-based paradigm that scores variants by evolutionary conservation -- an assumption that pharmacogenes including CYPs violate. Indeed, focusing on human CYPs, we show that homology-based models fail systematically. AlphaMissense (AM) assigns variants to its "ambiguous" class at nearly twice the proteome-wide rate across six CYPs, and within that class the scores are essentially uncorrelated with CYP2C9 DMS activity (Spearmans{rho} = 0.069); Evolutionary Scale Modeling 2 (ESM-2) shows the same pattern. We further hypothesized that non-homology-based features (sequence position, substitution chemistry, binding-site distance, and secondary structure) might help resolve the ambiguous calls, but their explanatory power is weak. Using CYP2C9 DMS activity as ground truth, we built a k-nearest-neighbors model over ESM-2 embeddings and ensembled it with AM and ESM-2 masked marginal probability, improving the ambiguous-class correlation roughly ten-fold, from{rho} = 0.069 to 0.715 (overall{rho} from 0.638 to 0.825). However, this markedly improved accuracy does not translate into agreement with clinical annotations. Drawing on evidence that a variants effect can depend on the drug, we hypothesize that substrate identity is the key missing feature in current models, and that predicting function for these multi-substrate enzymes may require redefining function as substrate-conditioned.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.