A curated reference dataset and deep learning model for multi-lead electrocardiographic interval measurements in UK Biobank
Kaplan, T.; Ramirez, J.; Young, W. J.; Sanghvi, M. M.; Madrid, J.; Naderi, H.; Shah, R.; Minchole, A.; Orini, M.; Tinker, A.; Lambiase, P. D.; Munroe, P. B.; van Duijvenboden, S.
Show abstract
Electrocardiographic (ECG) interval measurements underpin clinical decision-making and large-scale cardiovascular research, yet existing automated methods are often developed using small, heterogeneous datasets with limited expert annotation and uncertain generalisability to population cohorts. We developed a deep learning framework for automated PR, QRS, and QT interval estimation and established a large expert-curated reference dataset using UK Biobank (UKB) ECGs. The reference dataset comprises 11,330 lead-level annotations from 12-lead ECGs in 1,030 randomly selected UKB participants, generated using a standardised annotation protocol with independent expert review. A 1D convolutional neural network was trained to segment ECG waveforms and derive PR, QRS, and QT intervals. Performance was evaluated against expert annotations, UKB CardioSoft measurements, an open-source signal-processing toolbox, and a wavelet-based delineation method. Clinical validity was evaluated through associations with incident atrial fibrillation and major adverse cardiovascular events (MACE). Inter-observer agreement was high (ICC 0.81 - 0.97). In a held-out test set, the deep learning model achieved mean absolute errors of 7.7 ms (PR), 7.5 ms (QRS), and 4.9 ms (QT), outperforming all comparator methods, with minimal bias relative to expert annotations. Review of distributional outliers confirmed >80% validity for most interval measurements. Among 46,749 participants with follow-up (median 4 years), prolonged QTc derived by the deep learning model showed stronger associations with incident MACE (hazard ratio 2.9, 95% CI 2.1 - 4.0) than wavelet-based measurements (1.7, 1.4 - 2.0) or CardioSoft (1.1, 0.9 - 1.4). This expert-curated reference dataset and validated deep learning framework provide a scalable foundation for reproducible ECG phenotyping in UKB and beyond.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.