Long-term Penetrance of Disease Variants in Genes Prioritized for Genomic Newborn Screening: Evidence from Adult Biobanks
Gold, N. B.; Zouk, H.; Yeo, J.; Lipsitz, S.; Koyama, S.; Somanchi, H.; Perez, E.; Selvaraj, M. S.; O'Grady, L.; Miller, E.; Lewis, A. C. F.; Karlson, E. W.; Strong, A.; Gold, J. I.; Rehm, H. L.; Natarajan, P.; Green, R. C.
Show abstract
Importance: Genomic newborn screening (gNBS) is a potential public health intervention, but its positive predictive value (PPV) remains uncertain. Estimating the prevalence and penetrance of pathogenic and likely pathogenic (P/LP) variants in genes prioritized for screening may clarify the long-term PPV and clinical utility of gNBS. Objective: To compare ICD-based ascertainment, electronic medical record (EMR) review, and clinical assessment of genetic disorders in adults with P/LP variants in 54 genes prioritized for gNBS. Design: Two-cohort observational study with EMR review and clinical assessment in the hospital-based cohort. Setting: The U.K. Biobank (UKB) and Mass General Brigham Biobank (MGBB). Participants: 451,877 adults from the UKB and 53,371 from the MGBB, all with exome sequencing data. Exposures: P/LP variants in 54 genes prioritized through expert consensus for gNBS, in genotypes consistent with each gene's inheritance pattern. Main outcomes and measures: The primary outcome was the absolute difference in the proportion of MGBB participants identified as affected by ICD versus EMR ascertainment. Secondary outcomes included findings from clinical assessments of undiagnosed MGBB participants, corrected UKB penetrance estimates, and extrapolation to U.S.. annual birth cohorts and living adults. Results: P/LP variants were identified in 665 UKB participants (0.15%) and 82 MGBB participants (0.15%), approximately 1 in 650. In MGBB, EMR review revealed that 58/82 individuals (70.7%) were undiagnosed, although 25 of 58 (43.1%) had documented symptoms. Disease-associated ICD codes were found in 39.0% (32/82) of participants, whereas EMR review identified symptoms in 59.8% (49/82, McNemar P<.001). Applied to UKB, this correction yielded a penetrance of 28.4% (95% CI, 18.6% to 38.2%), implying that 73 to 203 participants beyond the 51 identified by ICD codes may have clinical features of disease. Extrapolated to U.S. birth cohorts, 4,900 to 5,700 newborns per year may harbor P/LP variants in these genes and survive into adulthood. Approximately 355,000 to 410,000 U.S. adults may have P/LP variants in these genes. Conclusions and relevance: Penetrance of P/LP variants in genes prioritized for gNBS is substantially higher than ICD estimates suggest. Many adults with P/LP variants are symptomatic but undiagnosed, supporting inclusion of these genes in gNBS.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Returning Actionable Genomic Results in a Research Biobank: Analytic Validity, Clinical Implementation and Resource Utilization 95%
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 95%
- Genome Sequencing and Comprehensive Rare Variant Analysis of 465 Families with Neurodevelopmental Disorders 94%
Similar papers in this journal
- Identification and validation of novel candidate risk genes in endocytic vesicular trafficking associated with esophageal atresia and tracheoesophageal fistulas 93%
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 92%
- A year of COVID-19 GWAS results from the GRASP portal reveals potential SARS-CoV-2 modifiers 91%
Similar papers in this journal
- Genetic Diagnosis of Facioscapulohumeral Muscular Dystrophy Type 1 Using Rare Variant Linkage Analysis and Long Read Genome Sequencing 93%
- A clinical algorithm to identify people with the glucose-6-phosphate dehydrogenase p.Val68Met variant at risk for diabetes undertreatment 92%
- 3-hour genome sequencing and targeted analysis to rapidly assess genetic risk 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.