Rigorous Female Breast Cancer Phenotyping Using the All of Us Research Program
Qi, Y.; Lundy-Perez, K.; Gee, D. A.; Chambwe, N.
Show abstract
Objectives Accurate phenotyping of cases and controls is essential for studying biological and environmental contributors to disease in large biobanks. We aimed to develop a flexible, customizable, and reproducible electronic health record (EHR)-based phenotyping framework for identifying disease cases and generating matched control cohorts for downstream analyses. Here, we developed the Phenotyping Algorithm for Cases and matched Controls using EHR-based Rules (PACER). Materials and Methods Applying PACER to the All of Us Research Program Curated Data Repository v8.0, we identified female breast cancer (BC) cases identified among participants recorded as female at birth using at least two BC-associated diagnostic Observational Medical Outcomes Partnership concept IDs documented at least 30 days apart. A one-to-one matched control cohort was generated by jointly matching on sex, age, genetic ancestry, and state-level residency. Clinical, socioeconomic, and genomic data were integrated for analysis. Results We identified 10,225 BC cases and generated a control cohort of the same size matched for key demographic characteristics. Comparison with a phecodeX-based BC cohort showed 91.03% agreement. Among cases responding to relevant survey items, 80.86% self-reported a personal history of BC, compared to 1.89% of controls. We detected an enrichment of BC-associated GWAS catalog variants, pathogenic mutations in known risk genes, and higher polygenic risk scores in cases compared to controls. Discussion and Conclusion Concordance across a phecodeX-based cohort, self-reported survey responses, and genomic analyses supports the validity of PACER-defined cohorts. PACER is publicly available and readily adaptable to other diseases, supporting future research in risk modeling and precision medicine.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The Multiethnic Cohort: A Resource for the study of Genetic and non-Genetic Cancer Risk Across Populations 93%
- Incorporating alternative Polygenic Risk Scores into the BOADICEA breast cancer risk prediction model 93%
- Causal effects of breast cancer risk factors across hormone receptor breast cancer subtypes: A two-sample Mendelian randomization study 92%
Similar papers in this journal
- Putative breast cancer risk variants from populations of South Asian ancestry are under-represented in public variant classification databases 94%
- A genome-wide association study of mammographic texture variation 93%
- Comparative validation of the BOADICEA and Tyrer-Cuzick breast cancer risk models incorporating classical risk factors and polygenic risk in a population-based prospective cohort 93%
Similar papers in this journal
Similar papers in this journal
- Investigating the relationship between breast cancer risk factors and an AI-generated mammographic texture feature in the Nurses' Health Study II 94%
- An updated PREDICT breast cancer prognostic model including the benefits and harms of radiotherapy 93%
- RNA Sequencing-Based Single Sample Predictors of Molecular Subtype and Risk of Recurrence for Clinical Assessment of Early-Stage Breast Cancer 93%
Similar papers in this journal
- Incorporating Polygenic Risk Scores and Nongenetic Risk Factors for Breast Cancer Risk Prediction among Asian Women, Results from Asia Breast Cancer Consortium 95%
- Missing data in the medical record for oncology patients: prevalence and association with outcomes 90%
- COVID-19 outcomes, risk factors and associations by race: a comprehensive analysis using electronic health records data in Michigan Medicine 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.