Polygenic scores in disease prediction: evaluation using the relevant performance metrics
Hingorani, A. D.; Gratton, J.; Finan, C.; Schmidt, A. F.; Patel, R.; Sofat, R.; Kuan, V.; Langenberg, C.; Hemingway, H.; Morris, J. K.; Wald, N. J.
Show abstract
BackgroundThe clinical value of polygenic risk scores has been questioned. We sought to clarify performance in population screening, individual risk prediction and population risk stratification by analysing 926 polygenic risk scores for 310 diseases from the Polygenic Score (PGS) Catalog. MethodsPolygenic risk scores in the PGS Catalog are reported using hazard ratios or odds ratios per standard deviation, or the area under the receiver operating characteristic curve sometimes expressed as the C-index. We used this information to produce estimates of performance in: (a) population screening -- by calculating the detection rate (DR5) for a 5% false positive rate (FPR) and the population odds of becoming affected given a positive result (OAPR); (b) individual risk prediction -- by calculating the individual odds of becoming affected for a person with a particular polygenic score; and (c) population risk stratification -- by calculating the odds of becoming affected for groups of individuals in different portions of a polygenic risk score distribution. We use coronary artery disease and breast cancer as illustrative examples. FindingsPopulation screening performance: The median DR5 for all polygenic risk scores and all diseases studied was 11% [interquartile range 8 - 18%]. The median DR5 was 12% [9 - 19] for polygenic risk scores for CAD and 10% [9 - 12] for breast cancer, with population OAPRs of 1:8 and 1: 21 respectively, with background 10-year odds of 1:19 and 1:41 respectively, which are typical for these diseases at age 50. Individual risk prediction: The corresponding 10-year odds of becoming affected for individuals aged 50 with a polygenic risk score at the 2.5th, 25th, 75th and 97.5th centile were 1:54, 1:29, 1:15, and 1:8 for CAD and 1:91, 1:56, 1:34, and 1:21 for breast cancer. Population risk stratification: At age 50, stratifying into quintile groups of CAD risk yielded 10-year odds of 1: 41 and 1: 11 for the lowest and highest quintile groups respectively. The 10-year odds was 1: 7 for the upper 2.5% of the polygenic risk score distribution for CAD, a group that contributed 7% of cases. The corresponding estimates for breast cancer were 1: 72 and 1: 26 for lowest and highest quintiles; and 1:19 for the upper 2.5% of the distribution, which contributed 6% of cases. InterpretationPolygenic risk scores perform poorly in population screening, individual risk prediction, and population risk stratification. FundingBritish Heart Foundation; UK Research and Innovation; National Institute of Health and Care Research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Disease-specific variant pathogenicity prediction significantly improves variant interpretation in inherited cardiac conditions 94%
- Leveraging pleiotropy to improve genetic risk prediction across diseases 93%
- Classification of Variants of Reduced Penetrance in High Penetrance Cancer Susceptibility Genes: Framework for Genetics Clinicians and Clinical Scientists by CanVIG-UK (Cancer Variant Interpretation Group-UK) 93%
Similar papers in this journal
- Genetic association studies using disease liabilities from deep neural networks 96%
- Widespread recessive effects on common diseases in a cohort of 44,000 British Pakistanis and Bangladeshis with high autozygosity 95%
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 95%
Similar papers in this journal
- A unified framework for estimating country-specific cumulative incidence for 18 diseases stratified by polygenic risk 95%
- Calibrated rare variant genetic risk scores for complex disease prediction using large exome sequence repositories 95%
- Transferability of genetic loci and polygenic scores for cardiometabolic traits in British Pakistanis and Bangladeshis 94%
Similar papers in this journal
- Augmenting clinical risk prediction of cardiovascular disease through protein and epigenetic biomarkers 94%
- Dynamic Importance of Genomic and Clinical Risk for Coronary Artery Disease Over the Life Course 94%
- A Multi-Ancestry Polygenic Risk Score for Coronary Heart Disease Based on an Ancestrally Diverse Genome-Wide Association Study and Population-Specific Optimization 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.