Handling onset age inconsistencies in longitudinal healthcare survey data
Li, W.; Yuan, M.; Park, Y.; Dao Duc, K.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWLongitudinal healthcare surveys frequently contain inconsistencies in self-reported onset ages, where participants report different ages for the same condition between enrollment and follow-up surveys. We propose two methods to handle this challenge. First, we introduce a procedure that aggregates inconsistency patterns to construct participant-level reliability scores, enabling researchers to stratify participants and prioritize analysis on high-reliability cohorts. Second, we present a Bayesian adjustment method that models enrollment and follow-up reports as noisy observations of a latent true onset age, producing adjusted estimates for the inconsistent observations that account for age-dependent and inter-survey-time effects. We evaluate both methods using data from the Canadian Partnership for Tomorrows Health. In general, both methods substantially strengthen correlations between biologically related conditions and improve predictive performance across classification and regression tasks. In addition, high-reliability cohorts from reliability score-based stratification reveal more coherent and interpretable disease clustering networks, and Bayesian adjustment shows particularly notable gains when multiple inconsistent variables are adjusted simultaneously. Finally, we provide guidance on choosing between these methods for healthcare practitioners. Institutional Review Board (IRB)The study is approved by the University of British Columbia IRB (IRB #H23-03800). Data and Code AvailabilityCanPath data are available to researchers through a controlled access process via the CanPath Access Portal (https://portal.canpath.ca). The code is available at https://anonymous.4open.science/r/canpath-FCCF.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using indication embeddings to represent patient health for drug safety studies 92%
- Modeling physician variability to prioritize relevant medical record information 91%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 91%
Similar papers in this journal
- Nearest-neighbor Projected-Distance Regression (NPDR) for detecting network interactions with adjustments for multiple tests and confounding 93%
- Joint Modeling of Longitudinal Biomarker and Survival Outcomes with the Presence of Competing Risk in Nested Case-Control Studies with Application to the TEDDY Microbiome Dataset 93%
- Robustifying genomic classifiers to batch effects via ensemble learning 93%
Similar papers in this journal
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 94%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 93%
- SynTL: A synthetic-data-based transfer learning approach for multi-center risk prediction 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.