Back

A Quantitative Framework for EHR Cohort Refinement.

Songthangtham, N.; Simon, G.; Johnson, S. G.

2026-04-28 health informatics
10.64898/2026.04.27.26351837 medRxiv
Show abstract

Electronic Health Record (EHR) based research depends on accurate cohort definitions, yet current workflows offer little early insight into cohort quality and rely heavily on slow, ad hoc manual review. We developed a quantitative, iterative framework that integrates model-guided case selection with sequential statistical testing to provide an early, data-driven signal of cohort accuracy. Using an OMOP-standardized dataset and the PCORnet Type 2 Diabetes phenotype as the gold-standard proxy, we evaluated eight sampling strategies across starting review sizes from 30 to 2,000 cases. At each iteration, the framework retrained an internal logistic regression model, selected cases for review, and applied a Bonferroni-adjusted Agresti-Coull upper bound to assess whether the cohort met a pre-specified accuracy threshold. Across 48 simulation conditions, starting sample size strongly shaped iteration count and total review burden, and adaptive sampling strategies consistently required fewer reviewed cases than fixed-batch methods. These findings demonstrate a reproducible, statistically grounded approach for refining EHR cohort definitions while reducing manual review effort.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.