The dangers of data double dipping in assessing the classification accuracies of blood biomarkers in Alzheimer's disease and related disorder research
Liu, T.; Zeng, X.; Snitz, B. E.; Karikari, T. K.; Deek, R. A.
Show abstract
Blood biomarker models are increasingly used in Alzheimer's disease and related dementia translational research, but predictive performance can be inflated when the same dataset is used for both model development and evaluation. We assess the effect of data double dipping using simulations and NULISA proteomic data from the MYHAT-NI community-based cohort to predict brain amyloid-beta neuroimaging status. In both settings, training AUC increased as more biomarkers were added, while testing AUC peaked earlier and then declined. These findings show that data double dipping can inflate model performance and highlight the need for external validation or internal validation with data partitioning.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable deep learning approach for extracting cognitive features from hand-drawn images of intersecting pentagons in older adults 93%
- An aging focused unobtrusive and Privacy-Preserving Digital Behaviorome 92%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 92%
Similar papers in this journal
- Targeted screening for Alzheimer’s disease clinical trials using data-driven disease progression models 94%
- An Explainable Multi-Modal Neural Network Architecture for Predicting Epilepsy Comorbidities Based on Administrative Claims Data 91%
- PECLIDES Neuro - A Personalisable Clinical Decision Support System for Neurological Diseases 90%
Similar papers in this journal
- AI reveals insights into link between CD33 and cognitive impairment in Alzheimer's Disease 94%
- A time-series analysis of blood-based biomarkers within a 25-year longitudinal dolphin cohort. 93%
- Bayesian Structural Time Series for Biomedical Sensor Data: A Flexible Modeling Framework for Evaluating Interventions 92%
Similar papers in this journal
- c-Triadem: A constrained, explainable deep learning model to identify novel biomarkers in Alzheimer’s disease 94%
- CohortDiagnostics: phenotype evaluation across a network of observational data sources using population-level characterization 93%
- Interpretable multivariate survival models: Improving predictions for conversion from mild cognitive impairment to Alzheimers disease (AD) via data fusion and machine learning 93%
Similar papers in this journal
- Identifying healthy individuals with Alzheimer neuroimaging phenotypes in the UK Biobank 94%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 92%
- Stratification of Alzheimer's Disease Patients Using Knowledge-Guided Unsupervised Latent Factor Clustering with Electronic Health Record Data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.