Back

When can predictive uncertainty be trusted? A methodological evaluation in free-living wearable electrocardiogram signal-quality assessment

Tran, K. D.

2026-08-28 health informatics
10.64898/2026.08.25.26361304 medRxiv
Show abstract

Uncertainty quantification is proposed as a safeguard for machine-learning systems in health-related signal analysis, but an uncertainty score is useful only if it behaves as a reliability signal. Free-living wearable electrocardiogram (ECG) signal-quality assessment provides a test bed because ambiguity, artifact, and acquisition shift can alter the relationship between confidence and correctness. This study evaluates predictive uncertainty under ambiguity, controlled corruption, and external distribution shift. 32,224 non-overlapping 10-s windows of synchronised single-lead ECG and three-axis accelerometry from 15 subjects in the Brno University of Technology ECG Quality Database were analysed. Two model families were compared: multinomial logistic regression and Classification and Regression Tree (CART), each progressing from a point estimate to a fixed-structure posterior and then a structure posterior. Expected conditional entropy and mutual information were evaluated as designated aleatoric and epistemic uncertainty measures, with max-softmax uncertainty as a confidence baseline. Validation covered error ranking, selective prediction, behavioural probes, posterior structural diversity, recorded-noise stress testing, and zero-shot external transfer. The logistic structure posterior retained an expected 8.5 of nine features and concentrated on near-complete masks, yielding little additional predictive diversity. Bayesian CART produced 221 distinct complete topologies among 238 retained draws and stronger score-dependent selective-risk behaviour. Conditional entropy increased with local class overlap, whereas mutual information increased when training information was reduced, although both showed cross-sensitivity. Under recorded noise, predicted quality severity changed more consistently than uncertainty, while external transfer preserved ordinal severity more reliably than uncertainty ordering. These findings show that posterior richness alone does not establish reliable uncertainty. Model-derived uncertainty should therefore be validated against prespecified ambiguity, information, and shift probes before supporting abstention, reacquisition, or downstream decisions.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.