Clinician Performance in Training Data Curation for an Arrhythmia Machine Learning Model: Is Anyone Qualified?
Kim, M.; Assadi, A.; Ehrmann, D.; Dixon, W.; Vecile, S.; Greer, R.; Goodfellow, S.; Bulic, A.
Show abstract
BackgroundEarly identification of arrhythmias in the intensive care unit (ICU) is important to prevent ICU morbidity and mortality. Timely arrhythmia detection relies on bedside providers telemetry interpretation. Machine learning (ML) models can function as clinical support tools to facilitate diagnoses. ML model development requires well-curated training data. The differential performance between labelers of different roles and experience is currently unknown. MethodsThis was a prospective observational study with frontline providers. 300 (200 original, 100 duplicate) 10-second telemetry tracings were labeled including sinus rhythm, 2nd/3rd degree atrioventricular (AV) block, junctional ectopic tachycardia (JET), ectopic atrial tachycardia (EAT), and reentrant supraventricular tachycardia (SVT). Interrater reliability was calculated against the ground truth label as the primary performance measure (intrarater reliability for consistency utilizing duplicate labels). Results11 participants completed the study: 1 Cardiology fellow, 4 pediatric ICU fellows, 2 pediatric cardiac ICU fellows, 3 pediatric cardiac ICU NPs, and 1 Pediatrics resident. Highest level of agreement was moderate ({kappa} 0.68, p <0.001) with the majority poor to moderate. There was no association of clinical subspecialty with labeling performance. Performance varied by rhythm type (median {kappa}): sinus (0.61), AV block (0.68) > Junctional (0.47), EAT (0.25), SVT (0.49). There was good intrarater reliability ({kappa} 0.71 [median], p<0.001). ConclusionsOverall, frontline provider performance was poor especially for complex arrhythmia classes. Cardiology training and experience was not associated with better performance. These findings highlight the need for thoughtful consideration in labeler training and validates the need for a clinical decision support tool in arrhythmia detection.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Early Experience with Ivabradine for Focal Atrial Tachycardia in Pediatric Patients with Congenital Heart Disease 94%
- Right bundle branch pacing: criteria, characteristics and outcomes 94%
- Fixation beats – a novel marker for reaching the left bundle branch area during deep septal lead implantation 93%
Similar papers in this journal
- Age Prediction From 12-lead Electrocardiograms Using Deep Learning: A Comparison of Four Models on a Contemporary, Freely Available Dataset 92%
- High-Fidelity Measurement of Pulse Arrival Time in Critically Ill Children Using Standard Bedside Monitoring Equipment 91%
- Classification of 12-lead ECGs: the PhysioNet/Computing in Cardiology Challenge 2020 91%
Similar papers in this journal
- QRS detection in single-lead, telehealth electrocardiogram signals: benchmarking open-source algorithms 96%
- Use of a Continuous Single Lead Electrocardiogram Analytic to Predict Patient Deterioration Requiring Rapid Response Team Activation 95%
- Detecting QT prolongation From a Single-lead ECG With Deep Learning 93%
Similar papers in this journal
- Applications of Machine Learning in Decision Analysis for Dose Management for Dofetilide 94%
- ChatGPT Provides Inconsistent Risk-Stratification of Patients With Atraumatic Chest Pain 93%
- Predictors and outcomes of Cardiac Dyssynchrony among patients with heart failure attending Benjamin Mkapa Hospital in Dodoma, central Tanzania: A protocol of prospective-longitudinal study 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.