Modeling diagnostic code dropout of schizophrenia in electronic health records improves phenotypic data quality and cross-ancestry transferability of polygenic scores
Burstein, D.; Tomasi, S.; Venkatesh, S.; VA Million Veteran Program, ; Rizk, M.; Roussos, P.; Voloudakis, G.
Show abstract
ImportanceResearchers commonly use counts of diagnostic codes from EHR-linked biobanks to infer phenotypic status. However, these approaches overlook temporal changes in EHR data, such as the discontinuation or "dropout" of diagnostic codes, which may exacerbate disparities in genomics research, as EHR data quality can be confounded with demographic attributes. ObjectiveTo address this, we propose modeling diagnostic code dropout in EHR data to inform phenotyping for schizophrenia in genomic analyses. DesignWe develop and test our diagnostic dropout model by analyzing EHR data from individuals with prior schizophrenia diagnoses. We further validate model performance on a subset of patients whose diagnoses were attained through chart review. Using PRS-CS and existing GWAS summary statistics, we first extrapolate polygenic weights. Then, we apply our dropout models outputs to construct a data-driven filter defining our target cohort for measuring polygenic score performance. SettingOur analysis utilizes EHR and genomic data from the Million Veteran Program. ParticipantsTo model diagnostic dropout in schizophrenia, we leverage data from 12,739 patients with a history of schizophrenia, after excluding outliers. For polygenic score analyses, we incorporate data from a potential pool of 8,385 European ancestry and 6,806 African ancestry patients with a history of schizophrenia. Main outcomes and measuresWe compare the performance of our diagnostic dropout model with alternative methodologies both in predicting diagnostic dropout on a holdout set, as well as on chart review labeled data. Using the top differential diagnosis predictors in our model, we select relevant cases by filtering out patients with a prior history of mood or anxiety disorders. We then test the impact of applying different filters for measuring polygenic score performance. ResultsWhen evaluated on chart review-labeled data, our model improves the area under the precision-recall curve (AUPRC) by 9.6% compared to competing methods. By applying our data-driven filter for schizophrenia, we achieve a 62% increase in the association effect size when transferring a European polygenic score to an African ancestry target cohort. Conclusions and RelevanceThese findings highlight the potential of modeling diagnostic code dropout to enhance the phenotypic quality of EHR-linked biobank data, advancing more equitable and accurate genomics research across diverse populations. Key PointsO_ST_ABSQuestionC_ST_ABSCan we leverage temporal changes in electronic health record (EHR) data to improve schizophrenia case selection for genomic studies? FindingsWe trained an XGBoost model on EHR data from 12,739 patients to predict schizophrenia diagnostic code dropout in the Million Veteran Program. By excluding cases with conditions associated with diagnostic dropout, we achieved a 62% increase in effect size when applying polygenic weights to an African ancestry target cohort. Filtering based on substance use, a common approach, yielded minimal gains. MeaningModeling diagnostic code dropout enhances the phenotypic quality of EHR-linked biobank data, and promotes equitable genomics research across diverse populations.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genetic implication of prenatal GABAergic and cholinergic neuron development in susceptibility to schizophrenia 96%
- A 10-Year Longitudinal Study of Brain Cortical Thickness in People with First-Episode Psychosis using Normative Models 94%
- Increased Prevalence of Rare Copy Number Variants in Treatment-Resistant Psychosis 94%
Similar papers in this journal
- Connectivity patterns of task-specific brain networks allow individual prediction of cognitive symptom dimension of schizophrenia and link to molecular architecture 92%
- Sex differences in the human brain transcriptome of cases with schizophrenia 92%
- Genomic stratification of clozapine prescription patterns using schizophrenia polygenic scores 92%
Similar papers in this journal
- Multivariate GWAS of psychiatric disorders and their cardinal symptoms reveal two dimensions of cross-cutting genetic liabilities 94%
- Systematic investigation of allelic regulatory activity of schizophrenia-associated common variants 92%
- The genetic and phenotypic correlates of neonatal Complement Component 3 and 4 protein concentrations with a focus on psychiatric and autoimmune disorders 91%
Similar papers in this journal
- Comparative genetic architectures of schizophrenia in East Asian and European populations 95%
- Genome-wide landscape of RNA-binding protein dysregulation reveals a major impact on psychiatric disorder risk 94%
- Identifying loci with different allele frequencies among cases of eight psychiatric disorders using CC-GWAS 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.