Calibrated Prediction Intervals for Polygenic Scores: Updated Comparisons, Contextual Calibration, and Data Normalization
Chang, X.; Hou, S.; Zhou, X.
Show abstract
Calibrated prediction intervals for polygenic scores (PGS) are essential for communicating individual-level uncertainty in genomic medicine. We present updated comparisons of two methods for constructing such intervals: CalPred, a parametric approach, and PredInterval, a non-parametric approach. Our results show that both methods can achieve calibrated coverage, although CalPred additionally requires a sufficiently large calibration set. The two methods also exhibit complementary trade-offs with respect to dataset size and risk identification. We further show that contextual calibration, as introduced in Hou et al. and followed in Shi et al., is most naturally achieved through appropriate phenotype normalization and data preprocessing. Apparent miscalibration can arise from inadequate normalization or from providing contextual information to some methods but not others. In UK Biobank, standard GWAS phenotype normalization procedures are sufficient to achieve contextual calibration for traits analyzed. In the extreme simulations of Hou et al. and Shi et al., supplying contextual covariates to PredInterval restores contextual calibration without normalization, and appropriate normalization can achieve contextual calibration without supplying covariates, while also substantially improving upstream tasks including association power and PGS accuracy. Together, these results underscore the central role of phenotype normalization and data preprocessing in GWAS analyses, including reliable uncertainty quantification for PGS.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Hidden structure in polygenic scores and the challenge of disentangling ancestry interactions in admixed populations 95%
- Robust, flexible, and scalable tests for Hardy-Weinberg Equilibrium across diverse ancestries 95%
- Hierarchical clustering of gene-level association statistics reveals shared and differential genetic architecture among traits in the UK Biobank 94%
Similar papers in this journal
- SparsePro: an efficient fine-mapping method integrating summary statistics and functional annotations 96%
- Leveraging expression from multiple tissues using sparse canonical correlation analysis (sCCA) and aggregate tests improves the power of transcriptome-wide association studies (TWAS) 96%
- Identifying Causal Variants by Fine Mapping Across Multiple Studies 96%
Similar papers in this journal
- LDAK-KVIK performs fast and powerful mixed-model association analysis of quantitative and binary phenotypes 97%
- Leveraging a machine learning derived surrogate phenotype to improve power for genome-wide association studies of partially missing phenotypes in population biobanks 96%
- Fast and flexible joint fine-mapping of multiple traits via the Sum of Single Effects model 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.