Overcome the Limitation of Phenome-Wide Association Studies (PheWAS): Extension of PheWAS to Efficient and Robust Large-Scale ICD Codes Analysis
Lin, Y.; Zhang, S.; Vessels, T. J.; Bastarache, L.; Bejan, C. A.; Hsi, R. S.; Phillips, E. J.; Ruderfer, D. M.; Pulley, J.; Edwards, T.; Wells, Q. S.; Warner, J. L.; Denny, J. C.; Roden, D. M.; Kang, H.; Xu, Y.
Show abstract
The Phenome-wide association studies (PheWAS) have become widely used for efficient, high-throughput evaluation of relationship between a genetic factor and a large number of disease phenotypes, typically extracted from a DNA biobank linked with electronic medical records (EMR). Phecodes, billing code-derived disease case-control status, are usually used as outcome variables in PheWAS and logistic regression has been the standard choice of analysis method. Since the clinical diagnoses in EMR are often inaccurate with errors which can lead to biases in the odds ratio estimates, much effort has been put to accurately define the cases and controls to ensure an accurate analysis. Specifically in order to correctly classify controls in the population, an exclusion criteria list for each Phecode was manually compiled to obtain unbiased odds ratios. However, the accuracy of the list cannot be guaranteed without extensive data curation process. The costly curation process limits the efficiency of large-scale analyses that take full advantage of all structured phenotypic information available in EMR. Here, we proposed to estimate relative risks (RR) instead. We first demonstrated the desired nature of RR that overcomes the inaccuracy in the controls via theoretical formula. With simulation and real data application, we further confirmed that RR is unbiased without compiling exclusion criteria lists. With RR as estimates, we are able to efficiently extend PheWAS to a larger-scale, phenome construction agnostic analysis of phenotypes, using ICD 9/10 codes, which preserve much more disease-related clinical information than Phecodes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Individual Reference Intervals for Personalized Interpretation of Clinical and Metabolomics Measurements 93%
- A Utility-Based Machine Learning-Driven Personalized Lifestyle Recommendation for Cardiovascular Disease Prevention 92%
- Computational Strategies in Nutrigenetics: Constructing a Reference Dataset of Nutrition-Associated Genetic Polymorphisms 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Body mass index and birth weight improve polygenic risk score for type 2 diabetes 91%
- Comprehensive profiling of genomic and transcriptomic differences between risk groups of lung adenocarcinoma and lung squamous cell carcinoma 91%
- Development and validation of decision rules models to stratify coronary artery disease, diabetes, and hypertension risk in preventive care: cohort study of returning UK Biobank participants 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.