Gene identity, not variant effect, dominates ClinVar benchmarks of missense pathogenicity predictors
Arnoult, D.; Harrizi, S.; Nait Irahal, I.; Mostafa, K.
Show abstract
Missense pathogenicity predictors are routinely benchmarked against ClinVar, whose labels are strongly structured by gene: genes under diagnostic scrutiny accumulate pathogenic submissions while incidentally sequenced genes accumulate benign ones. We asked how much of a benchmark score this structure alone can produce. On 197,904 ClinVar missense variants validated against UniProt canonical sequences, a null model using no variant-level information, scoring each variant only by the pathogenic fraction of its own gene, reaches an area under the receiver operating characteristic curve (AUROC) of 0.921 under a random 10-fold split. On a common intersection of 169,989 variants, four current predictors exceed it by only 0.036 to 0.044. The inflation is not uniform, so it does not cancel when predictors are compared: under within-gene evaluation the ranking inverts, AlphaMissense rising from third to first and gMVP falling to third (p < 0.0001). The inversion survives removal of ceiling genes and replicates on an independently curated benchmark. Because both rankings derive from the same ClinVar labels, we arbitrated between them using data with no gene-level structure: agreement with 47 human deep mutational scanning assays matches the within-gene ranking and inverts the conventional one (p = 0.027, 0.0023). Across twenty-two dbNSFP predictors scored on one common intersection of 112,248 variants, with each tool's exposure to clinical labels registered before any score was extracted, predictors never trained on such labels sit 0.051 AUROC behind supervised ones globally but only 0.026 behind within genes (difference +0.025 [+0.023, +0.027], p < 0.0001). Leave-one-out correction, the standard remedy, is worth 0.002 AUROC. Much of ClinVar benchmark performance reflects gene identity rather than variant effect, and the distortion changes which predictor a benchmark ranks first, in a direction experimental data contradicts. We release genenull, a single-file implementation, so reporting this baseline costs one function call.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Juggling offsets unlocks RNA-seq tools for fast scalable differential usage, aberrant splicing and expression analyses. 92%
- Variant effect predictor correlation with functional assays is reflective of clinical classification performance 92%
- Empowering rare variant burden-based gene-trait association studies via optimized computational predictor choice 92%
Similar papers in this journal
- Evidence-based calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for clinical use of PP3/BP4 criteria 92%
- GA4GH Phenopacket-Driven Characterization of Genotype-Phenotype Correlations in Mendelian Disorders 92%
- Availability of benign missense variant “truthsets” for validation of functional assays: current status and a novel systematic approach 91%
Similar papers in this journal
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 94%
- RNAIndel: a machine-learning framework for discovery of somatic coding indels using tumor RNA-Seq data 92%
- AutoPM3: Enhancing Variant Interpretation via LLM-driven PM3 Evidence Extraction from Scientific Literature 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.