Back

Pervasive ancestry bias in variant effect predictors

Pathak, A. K.; Bora, N.; Badonyi, M.; Livesey, B. J.; SG10K_Health Consortium, ; Ngeow, J.; Marsh, J. A.

2025-03-25 bioinformatics
10.1101/2024.05.20.594987 bioRxiv
Show abstract

Variant effect predictors (VEPs) - computational tools that assess the potential impact of genetic variants - have become increasingly vital for clinical variant interpretation. Currently, most VEPs used in variant prioritisation have been trained on datasets of clinically curated or population-derived variants. These datasets, however, disproportionately represent individuals of European descent. We hypothesised that this bias may lead to unequal VEP performance across different populations. To test this, we evaluated the scoring patterns of 52 VEPs for missense variants across 14 ancestry groups. We observe striking disparities: some VEPs predict a markedly higher proportion of damaging variants in underrepresented populations, such as those of Malay descent, compared to individuals of European ancestry. In contrast, VEPs that do not rely on clinical or population data predict more consistent pathogenicity burdens across ancestry groups. Moreover, we could closely link these discrepancies across methods to biases in training data. Our findings underscore the urgent need to adopt tools that minimise ancestry bias to ensure fairer and more accurate variant effect prediction and genetic diagnoses for all populations.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.