Assessing the Accuracy of CGM Metrics: The Role of Missing Data and Imputation Strategies
Cichosz, S. L.; Kronborg, T.; Hangaard, S.; Vestergaard, P.; Jensen, M. H.
Show abstract
AimThis study aims to evaluate the accuracy of continuous glucose monitoring (CGM)-derived metrics, particularly those related to glycemic variability, in the presence of missing data. It systematically examines the effects of different missing data patterns and imputation strategies on both standard glycemic metrics and complex variability metrics. MethodsThe analysis modeled and compared the effects of three types of missing data patterns--missing completely at random (MCAR), segmental gaps, and blockwise gaps--with proportions ranging from 5% to 50% on CGM metrics derived from 14-day profiles of individuals with type 1 and type 2 diabetes. Six imputation strategies were assessed: data removal, linear interpolation, mean imputation, piecewise cubic Hermite interpolation, temporal alignment imputation (TAI) and random forest-based imputation. ResultsA total of 933 14-day CGM profiles from 468 individuals with diabeteswere analyzed. Across all metrics, the coefficient of determination (R2) improved as the proportion of missing data decreased, regardless of the missing data pattern. The impact of missing data on the agreement between imputed and reference metrics varied depending on the missing data pattern. To achieve high accuracy (R2 > 0.95) in representing true metrics, at least 80% of the CGM data was required. While no imputation strategy fully compensated for high levels of missing data, simple removal and TAI outperformed others in certain scenarios. ConclusionThis study highlights the significant impact of missing data and imputation strategies on CGM-derived metrics, particularly glycemic variability and time below range (TBR) estimates. The findings underscore the necessity of rigorous data handling practices to ensure reliable assessments of glycemic control and variability.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 95%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 92%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 90%
Similar papers in this journal
- Derivation and validation of a machine learning risk score using biomarker and electronic patient data to predict rapid progression of diabetic kidney disease 92%
- Replication and cross-validation of T2D subtypes based on clinical variables: an IMI-RHAPSODY study 92%
- Subgroups of young type 2 diabetes in India reveal insulin deficiency as a major driver 91%
Similar papers in this journal
- Leveraging Large Language Models to Analyze Continuous Glucose Monitoring Data: A Case Study 95%
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 92%
- Testing the phenotypic decanalization hypothesis: social determinants of hyperglycemia and type 2 diabetes in adult urban Argentinian population 92%
Similar papers in this journal
- Clinical characteristics and outcomes in diabetes patients admitted with COVID-19 in Dubai: a cross-sectional single centre study. 91%
- A Multivariate Forecasting Model for the COVID-19 Hospital Census Based on Local Infection Incidence 89%
- Uncovering clinical risk factors and prediction of severe COVID-19: A machine learning approach based on UK Biobank data 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.