Back

Combining MAVEs and computational predictors improves variant classification across ancestries in hereditary cancer genes

Bora, N.; Badonyi, M.; Pathak, A. K.; Manisekaran, R. G.; SG10K_Health Consortium, ; Marsh, J. A.; Ngeow, J.

2025-12-09 genetic and genomic medicine
10.64898/2025.12.08.25341119 medRxiv
Show abstract

Many commonly used computational tools for variant effect prediction exhibit ancestry-related bias because they are trained on clinical or population datasets that under-represent global diversity, leading to uneven and sometimes unfair variant classification across ancestries. Multiplexed assays of variant effect (MAVEs) and population-free VEPs instead offer alternatives that are unbiased with respect to human ancestry, providing classification evidence that generalises across populations. Here, we evaluate MAVE- and VEP-based classification across five cancer-associated genes with high-quality MAVE datasets, focusing on three leading population-free VEPs: GEMME, EVE, and CPT-1. We find that MAVEs are more conservative and decisive in classification, assigning fewer variants to the pathogenic category while yielding fewer indeterminate classifications when calibrated to ACMG/AMP guidelines. While MAVEs show heightened sensitivity in functionally assayed regions, VEPs identify a broader range of pathogenic variants overall. By combining clinical evidence strengths from MAVEs and VEPs, we reclassify over 90% of variants of uncertain significance across the SG10K_Health and Mexico City Prospective Study reference datasets as at least likely benign or likely pathogenic. We further isolate variants where MAVE- and VEP-based classifications are discordant, highlighting method-specific limitations. Together, these findings clarify the complementary strengths of experimental and computational classification approaches and provide a path to less biased and more equitable variant interpretation in clinical genomics, helping to mitigate disparities in diagnosis across ancestries.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.