Back

Reducing misclassification bias in electronic health record-based GWAS of psychiatric traits using the SuperControl framework

Eick, L.; FinnGen, ; ganna, a.; Yang, Z.

2025-12-15 genetic and genomic medicine
10.64898/2025.12.14.25342213 medRxiv
Show abstract

Electronic health records (EHRs) have enabled large-scale genetic studies of psychiatric disorders, but their diagnostic imprecision introduces substantial misclassification bias. This challenge is particularly pronounced for psychiatric traits, which lack objective biomarkers, exhibit high comorbidity, and are often underdiagnosed due to stigma and help-seeking barriers. We first used simulations to examine how misclassification bias interacts with the classic "super-normal" control design. The results showed that, under moderate to high misclassification, especially for more prevalent traits, excluding individuals with psychiatric comorbidities from the controls can recover power and reduce false positives. Guided by these insights, we constructed a refined SuperControl phenotype in FinnGen for ten psychiatric disorders. Despite reduced sample size, this approach increased genome-wide locus discovery and enhanced biological specificity, with stronger CNS heritability enrichment and improved cross-biobank polygenic prediction, while modestly increasing genetic correlations between traits. We further developed PRISMA, a machine-learning approach that complements the SuperControl phenotype by leveraging excluded individuals through imputation of disease liability. PRISMA further increased locus discovery, revealing significant associations for previously underpowered traits, though with greater pleiotropy and reduced tissue specificity. Together, these findings demonstrate that misclassification in EHR-derived psychiatric phenotypes can meaningfully suppress genetic signal, and that structured control refinement can mitigate this bias. Our results suggest that stricter cohort selection, supplemented by predictive imputation when appropriate, offers a scalable strategy to enhance discovery and generalizability in EHR-based psychiatric genetics.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.