Back

A curated reference dataset and deep learning model for multi-lead electrocardiographic interval measurements in UK Biobank

Kaplan, T.; Ramirez, J.; Young, W. J.; Sanghvi, M. M.; Madrid, J.; Naderi, H.; Shah, R.; Minchole, A.; Orini, M.; Tinker, A.; Lambiase, P. D.; Munroe, P. B.; van Duijvenboden, S.

2026-06-29 cardiovascular medicine
10.64898/2026.06.26.26356665 medRxiv
Show abstract

Electrocardiographic (ECG) interval measurements underpin clinical decision-making and large-scale cardiovascular research, yet existing automated methods are often developed using small, heterogeneous datasets with limited expert annotation and uncertain generalisability to population cohorts. We developed a deep learning framework for automated PR, QRS, and QT interval estimation and established a large expert-curated reference dataset using UK Biobank (UKB) ECGs. The reference dataset comprises 11,330 lead-level annotations from 12-lead ECGs in 1,030 randomly selected UKB participants, generated using a standardised annotation protocol with independent expert review. A 1D convolutional neural network was trained to segment ECG waveforms and derive PR, QRS, and QT intervals. Performance was evaluated against expert annotations, UKB CardioSoft measurements, an open-source signal-processing toolbox, and a wavelet-based delineation method. Clinical validity was evaluated through associations with incident atrial fibrillation and major adverse cardiovascular events (MACE). Inter-observer agreement was high (ICC 0.81 - 0.97). In a held-out test set, the deep learning model achieved mean absolute errors of 7.7 ms (PR), 7.5 ms (QRS), and 4.9 ms (QT), outperforming all comparator methods, with minimal bias relative to expert annotations. Review of distributional outliers confirmed >80% validity for most interval measurements. Among 46,749 participants with follow-up (median 4 years), prolonged QTc derived by the deep learning model showed stronger associations with incident MACE (hazard ratio 2.9, 95% CI 2.1 - 4.0) than wavelet-based measurements (1.7, 1.4 - 2.0) or CardioSoft (1.1, 0.9 - 1.4). This expert-curated reference dataset and validated deep learning framework provide a scalable foundation for reproducible ECG phenotyping in UKB and beyond.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
European Heart Journal - Digital Health
18 papers in training set
Top 0.1%
38.0%
2
Circulation: Genomic and Precision Medicine
48 papers in training set
Top 0.1%
9.3%
3
European Heart Journal
22 papers in training set
Top 0.3%
5.3%
50% of probability mass above
4
Heart
11 papers in training set
Top 0.1%
4.7%
5
Frontiers in Cardiovascular Medicine
53 papers in training set
Top 0.7%
4.2%
6
Circulation
74 papers in training set
Top 0.8%
3.9%
7
Journal of the American Heart Association
140 papers in training set
Top 2%
3.3%
8
Nature Communications
5641 papers in training set
Top 37%
3.1%
9
BMC Cardiovascular Disorders
18 papers in training set
Top 0.3%
3.0%
10
PLOS ONE
5266 papers in training set
Top 46%
2.0%
11
JACC: Clinical Electrophysiology
13 papers in training set
Top 0.1%
1.9%
12
The American Journal of Cardiology
17 papers in training set
Top 0.8%
1.6%
13
Scientific Reports
3612 papers in training set
Top 60%
1.4%
14
Heart Rhythm
23 papers in training set
Top 0.5%
1.1%
15
Open Heart
21 papers in training set
Top 0.9%
1.1%
16
Journal of the American College of Cardiology
12 papers in training set
Top 0.5%
1.1%
17
eLife
5828 papers in training set
Top 59%
1.1%
18
Nature Medicine
125 papers in training set
Top 4%
0.6%
19
Biology Methods and Protocols
61 papers in training set
Top 3%
0.6%
20
BMJ Open
601 papers in training set
Top 14%
0.6%
21
eBioMedicine
183 papers in training set
Top 8%
0.6%
22
PLOS Computational Biology
1863 papers in training set
Top 22%
0.6%
23
npj Digital Medicine
118 papers in training set
Top 4%
0.6%