Trajectories matter: Discovery and validation of ordered EHR sequences that inform clinical risk predictions
Edelman, B.; Kim, H.; Skolnick, J.
Show abstract
Structured AbstractO_ST_ABSObjectiveC_ST_ABSTo test whether temporally ordered clinical sequences mined from EHRs improve the understanding of downstream adverse outcomes versus unordered bag-of-codes representations. Materials and MethodsWithin the NIH All of Us Controlled Tier, we mined frequent ordered event pairs (A[->]B) and tested for risk elevation of later adverse events (C) to create three-event (A[->]B[->]C) sequences. The algorithm used observation-period-aware indexing, a maximum 5-year A[->]B gap, and 90-day latency before follow-up. Comparators were (B without prior A)[->]C (primary), B[->]A[->]C, (A without B)[->]C, and a calendar-time baseline. Covariates were balanced with inverse probability of treatment weighting (IPTW) Weighted Aalen-Johansen estimated cumulative incidence at 1, 2, and 5 years. Discovery and confirmation were analyzed separately with global Benjamini- Hochberg false discover rate (BH-FDR). ResultsOf 633,545 persons, 432,617 met eligibility ([≥]365 observed days) and were split 70/30 into discovery and confirmation. We mined 3,066,183 trajectories from the discovery set by combining 340,687 sufficiently supported A[->]B pairs with each of 9 curated adverse third events. Due to compute constraints, we tested 20,565 of these trajectories (0.67%) for risk elevation. After discovery FDR, 234 trajectories advanced, and 39 validated for at least one time horizon in confirmation. At 5 years, the median risk ratio (RR) was 3.96 versus baseline and 2.18 versus B[->]C with no prior A. Reverse-order checks were feasible for 89.7% of hypotheses; the median (A[->]B[->]C) vs (B[->]A[->]C) RR was 1.21. DiscussionOrdered trajectories captured clinically coherent pathways where temporal order added information beyond diagnosis presence alone. ConclusionTrajectory mining and confirmation reveal actionable, risk pathways that complement conventional risk models.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 95%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 95%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 95%
Similar papers in this journal
Similar papers in this journal
- A Neanderthal OAS1 isoform Protects Against COVID-19 Susceptibility and Severity: Results from Mendelian Randomization and Case-Control Studies 93%
- Actionable druggable genome-wide Mendelian randomization identifies repurposing opportunities for COVID-19 92%
- Exome-by-phenome-wide rare variant gene burden association with electronic health record phenotypes 92%
Similar papers in this journal
- A systematic analysis of the contribution of genetics to multimorbidity and comparisons with primary care data 93%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 92%
- Association between circulating inflammatory markers and adult cancer risk: a Mendelian randomization analysis 92%
Similar papers in this journal
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 91%
- Developing and Evaluating Pediatric Phecodes (Peds-Phecodes) for High-Throughput Phenotyping Using Electronic Health Records 91%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.