Back

Digitisation and linkage of PDF formatted 12-lead ECGs in Adult Congenital Heart Disease

Alkan, M.; Deligianni, F.; Anagnostopoulos, C.; Zakariyya, I.; Veldtman, G.

2024-12-16 health informatics
10.1101/2024.12.16.24319092 medRxiv
Show abstract

BACKGROUND12-lead ECGs form an essential part of the late follow-up of adults with congenital heart disease (ACHD). Such ECGs are most frequently reviewed by clinicians in paper or PDF formats. These visual representations of the original vector data do not easily lend themselves to be directly analysed with the increasingly powerful Machine Learning algorithms that hold promise in risk prediction and early prevention of adverse events. OBJECTIVESIn this work, we set out to recreate the original digital signals from ECG PDF documents by a series of data processing steps, validate accuracy of the process, and demonstrate its potential utility in research. METHODSUsing 4153 ECG PDF documents from 436 ACHD patients, we created a "pipeline" to successfully digitise the visually represented ECG vector datasets. We then proceed with the validation of the digitised ECG dataset using several features that are also calculated by the vendor, such as QRS duration, PR interval and ventricular rate, on all the patients. RESULTSWe confirmed a strong correlation with the vendor measured ECG parameters including PR interval (R = 0.941, P < 0.05), QRS duration (R = 0.949, P < 0.05) and ventricular rate (R = 0.971, P < 0.05). Further, using Support Vector Machine (SVM), a well-established Machine Learning (ML) model we demonstrate the ability of the digitised ECG dataset to accurately predict anatomic diagnosis in ACHD. CONCLUSIONSDigitisation of PDF formatted ECG signal data can be accomplished with good accuracy and can be used in clinical research in ACHD.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.