A Quantitative Framework for EHR Cohort Refinement.
Songthangtham, N.; Simon, G.; Johnson, S. G.
Show abstract
Electronic Health Record (EHR) based research depends on accurate cohort definitions, yet current workflows offer little early insight into cohort quality and rely heavily on slow, ad hoc manual review. We developed a quantitative, iterative framework that integrates model-guided case selection with sequential statistical testing to provide an early, data-driven signal of cohort accuracy. Using an OMOP-standardized dataset and the PCORnet Type 2 Diabetes phenotype as the gold-standard proxy, we evaluated eight sampling strategies across starting review sizes from 30 to 2,000 cases. At each iteration, the framework retrained an internal logistic regression model, selected cases for review, and applied a Bonferroni-adjusted Agresti-Coull upper bound to assess whether the cohort met a pre-specified accuracy threshold. Across 48 simulation conditions, starting sample size strongly shaped iteration count and total review burden, and adaptive sampling strategies consistently required fewer reviewed cases than fixed-batch methods. These findings demonstrate a reproducible, statistically grounded approach for refining EHR cohort definitions while reducing manual review effort.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 94%
- Using computable knowledge mined from the literature to elucidate confounders for EHR-based pharmacovigilance 94%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 93%
Similar papers in this journal
- Scalable information extraction from free text electronic health records using large language models 93%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 92%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 92%
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 93%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 92%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.