When can predictive uncertainty be trusted? A methodological evaluation in free-living wearable electrocardiogram signal-quality assessment
Tran, K. D.
Show abstract
Uncertainty quantification is proposed as a safeguard for machine-learning systems in health-related signal analysis, but an uncertainty score is useful only if it behaves as a reliability signal. Free-living wearable electrocardiogram (ECG) signal-quality assessment provides a test bed because ambiguity, artifact, and acquisition shift can alter the relationship between confidence and correctness. This study evaluates predictive uncertainty under ambiguity, controlled corruption, and external distribution shift. 32,224 non-overlapping 10-s windows of synchronised single-lead ECG and three-axis accelerometry from 15 subjects in the Brno University of Technology ECG Quality Database were analysed. Two model families were compared: multinomial logistic regression and Classification and Regression Tree (CART), each progressing from a point estimate to a fixed-structure posterior and then a structure posterior. Expected conditional entropy and mutual information were evaluated as designated aleatoric and epistemic uncertainty measures, with max-softmax uncertainty as a confidence baseline. Validation covered error ranking, selective prediction, behavioural probes, posterior structural diversity, recorded-noise stress testing, and zero-shot external transfer. The logistic structure posterior retained an expected 8.5 of nine features and concentrated on near-complete masks, yielding little additional predictive diversity. Bayesian CART produced 221 distinct complete topologies among 238 retained draws and stronger score-dependent selective-risk behaviour. Conditional entropy increased with local class overlap, whereas mutual information increased when training information was reduced, although both showed cross-sensitivity. Under recorded noise, predicted quality severity changed more consistently than uncertainty, while external transfer preserved ordinal severity more reliably than uncertainty ordering. These findings show that posterior richness alone does not establish reliable uncertainty. Model-derived uncertainty should therefore be validated against prespecified ambiguity, information, and shift probes before supporting abstention, reacquisition, or downstream decisions.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparison of wearable and clinical devices for acquisition of peripheral nervous system signals 94%
- On the reliability of wearable technology: A tutorial on measuring heart rate and heart rate variability in the wild 91%
- Feasibility of Ultra-Short Term Analysis of Heart Rate and Systolic Arterial Pressure Variability at Rest and During Stress via Time-domain and Entropy-based Measures 91%
Similar papers in this journal
- Effects of Electrocardiograms QRS Detection Algorithms in Heart Rate Variability Metrics 94%
- Explaining Deep Neural Networks for Knowledge Discovery in Electrocardiogram Analysis 93%
- An infection prediction model developed from inpatient data can predict out-of-hospital COVID-19 infections from wearable data when controlled for dataset shift 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Rett syndrome severity estimation with the BioStamp nPoint using interactions between heart rate variability and body movement 94%
- Discrimination of sleep and wake periods from a hip-worn raw acceleration sensor using recurrent neural networks 92%
- Applying time series analyses on continuous accelerometry data: a clinical example in older adults with and without cognitive impairment 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.