Genomic heterogeneity inflates the performance of variant pathogenicity predictions
Lu, B.; Liu, X.; Lin, P.-Y.; Brandes, N.
Show abstract
Recent studies have reported unprecedented accuracy predicting pathogenic variants across the genome, including in noncoding regions, using large AI models trained on vast genomic data. We present a comprehensive evaluation of these frontier models, showing that performance is inflated by differences in the prevalence of pathogenic variants across genomic contexts. We identify the best-performing models for each variant type and establish a benchmark to guide future progress.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MINTIE: identifying novel structural and splice variants in transcriptomes using RNA-seq data 95%
- Empowering rare variant burden-based gene-trait association studies via optimized computational predictor choice 95%
- Samplot: A Platform for Structural Variant Visual Validation and Automated Filtering 95%
Similar papers in this journal
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 96%
- Detection of aberrant splicing events in RNA-seq data with FRASER 96%
- G4mer: An RNA language model for transcriptome-wide identification of G-quadruplexes and disease variants from population-scale genetic data 95%
Similar papers in this journal
- Omics-informed CNV calls reduce false positive rate and improve power for CNV-trait associations 93%
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 93%
- Leveraging TOPMed Imputation Server and Constructing a Cohort-Specific Imputation Reference Panel to Enhance Genotype Imputation among Cystic Fibrosis Patients 93%
Similar papers in this journal
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 96%
- AutoPM3: Enhancing Variant Interpretation via LLM-driven PM3 Evidence Extraction from Scientific Literature 94%
- Estimating the power of sequence covariation for detecting conserved RNA structure 93%
Similar papers in this journal
- Advanced variant classification framework reduces the false positive rate of predicted loss of function (pLoF) variants in population sequencing data 96%
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 95%
- Evidence-based recommendations for gene-specific ACMG/AMP variant classification from the ClinGen ENIGMA BRCA1 and BRCA2 Variant Curation Expert Panel 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.