Accurate overall, uneven by patient: a benchmark and demographic audit of deep learning for 12 lead ECG classification on PTB-XL
Rehman, A. D.; Nazir, S.
Show abstract
Deep learning reads 12 lead electrocardiograms at close to expert level on public benchmarks, yet most reports give one accuracy figure for the whole test set and stop there. We trained three architectures that are standard in this field, a 1D ResNet, a convolutional network with a bidirectional LSTM, and a convolutional network with a bidirectional LSTM followed by a transformer encoder, on the PTB-XL dataset to classify the five diagnostic superclasses, and then looked at how each one performed across sex and age. On the held out fold all three reached a macro AUC near 0.92, in line with the strongest published results on this benchmark, and the simplest model, the 1D ResNet, was marginally the best at 0.9241. The averages hid a steady pattern. Every model scored lower for female patients than for male patients, and every model scored lowest for patients aged 80 and over, where the 1D ResNet fell to 0.8878 and the transformer to 0.8693. Adding complexity did not close either gap and slightly widened the gap by age. Overall accuracy on PTB-XL is close to solved for these model families, but the benefit is not shared evenly, and a single headline number hides the patients a model serves worst. We release the full stratified evaluation to support fairness aware reporting.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multiple Instance Learning Framework can Facilitate Explainability in Murmur Detection 95%
- A recurrent neural network and parallel hidden Markov model algorithm to segment and detect heart murmurs in phonocardiograms 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 93%
Similar papers in this journal
- Age Prediction From 12-lead Electrocardiograms Using Deep Learning: A Comparison of Four Models on a Contemporary, Freely Available Dataset 94%
- Classification of 12-lead ECGs: the PhysioNet/Computing in Cardiology Challenge 2020 93%
- Prospective validation of clinical deterioration predictive models prior to intensive care unit transfer among patients admitted to acute care cardiology wards 90%
Similar papers in this journal
- Biometric Contrastive Learning for Data-Efficient Deep Learning from Electrocardiographic Images 95%
- A Comparative Analysis of Privacy-Preserving Large Language Models For Automated Echocardiography Report Analysis 92%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 91%
Similar papers in this journal
- Rett syndrome severity estimation with the BioStamp nPoint using interactions between heart rate variability and body movement 92%
- Enhanced machine learning and hybrid ensemble approaches for coronary heart disease prediction 92%
- DeepGANnel: Synthesis of fully annotated single molecule patch-clamp data using generative adversarial networks 91%
Similar papers in this journal
- BRAVEHEART: Open-source software for automated electrocardiographic and vectorcardiographic analysis 93%
- Digitizing ECG image: new fully automated method and open-source software code 92%
- A standardized analytics pipeline for reliable and rapid development and validation of prediction models using observational health data 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.