Back

Clinician Performance in Training Data Curation for an Arrhythmia Machine Learning Model: Is Anyone Qualified?

Kim, M.; Assadi, A.; Ehrmann, D.; Dixon, W.; Vecile, S.; Greer, R.; Goodfellow, S.; Bulic, A.

2026-01-08 cardiovascular medicine
10.64898/2026.01.07.26343629 medRxiv
Show abstract

BackgroundEarly identification of arrhythmias in the intensive care unit (ICU) is important to prevent ICU morbidity and mortality. Timely arrhythmia detection relies on bedside providers telemetry interpretation. Machine learning (ML) models can function as clinical support tools to facilitate diagnoses. ML model development requires well-curated training data. The differential performance between labelers of different roles and experience is currently unknown. MethodsThis was a prospective observational study with frontline providers. 300 (200 original, 100 duplicate) 10-second telemetry tracings were labeled including sinus rhythm, 2nd/3rd degree atrioventricular (AV) block, junctional ectopic tachycardia (JET), ectopic atrial tachycardia (EAT), and reentrant supraventricular tachycardia (SVT). Interrater reliability was calculated against the ground truth label as the primary performance measure (intrarater reliability for consistency utilizing duplicate labels). Results11 participants completed the study: 1 Cardiology fellow, 4 pediatric ICU fellows, 2 pediatric cardiac ICU fellows, 3 pediatric cardiac ICU NPs, and 1 Pediatrics resident. Highest level of agreement was moderate ({kappa} 0.68, p <0.001) with the majority poor to moderate. There was no association of clinical subspecialty with labeling performance. Performance varied by rhythm type (median {kappa}): sinus (0.61), AV block (0.68) > Junctional (0.47), EAT (0.25), SVT (0.49). There was good intrarater reliability ({kappa} 0.71 [median], p<0.001). ConclusionsOverall, frontline provider performance was poor especially for complex arrhythmia classes. Cardiology training and experience was not associated with better performance. These findings highlight the need for thoughtful consideration in labeler training and validates the need for a clinical decision support tool in arrhythmia detection.

Published in CJC Pediatric and Congenital Heart Disease · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.