Detecting Cardiac Amyloidosis in Italian Cardiology Reports: Structured Variable Extraction versus Direct Free-Text Analysis
Mazzucato, S.; Sartiano, D.; Vergaro, G.; Dalmiani, S.; Emdin, M.; Micera, S.; Oddo, C. M.; Passino, C.; Moccia, S.; Bandini, A.; Seinen, T.
Show abstract
BackgroundEarly and accurate identification of cardiac amyloidosis improves patient outcomes, yet relevant evidence is frequently hidden in free-text records. This study assesses whether structured variable extraction or direct free-text analysis more reliably identifies patients with cardiac amyloidosis, with the goal of informing clinical decision support strategies. MethodsWe extracted 21 clinical variables from 432 Italian patient records using supervised and prompt-based methods with both proprietary and locally-deployable computational models. Classification performance was evaluated by comparing extracted data with gold-standard manual annotations. Two feature sets were tested: general clinical variables and amyloidosis-specific risk factors. Additionally, we evaluated direct zero-shot prediction on unstructured clinical notes. ResultsFor entity extraction, GPT-4.1-mini achieved F1=0.96, comparable to supervised SpaCy (F1=0.95) and GPT-4o (F1=0.94). Local open-source models like Qwen2.5 reached F1=0.94. For cardiac amyloidosis classification, machine learning models using full extracted features matched gold-standard results (SauerkrautLM-Gemma: F1=0.80 vs. gold: F1=0.82). General clinical features alone yielded lower performance (F1=0.68), highlighting that amyloidosis-specific risk factors in unstructured text provide discriminative diagnostic value. Zero-shot direct predictions outperformed supervised feature-based approaches (MedGemma: F1=0.92). ConclusionsAutomated extraction and zero-shot prediction effectively structure Italian EHRs and identify amyloidosis patients without manual annotation. Domain-specific risk factors in free-text notes provide substantial predictive value. Italian hospitals can potentially deploy locally-deployable models to screen cardiac amyloidosis without manual annotation or proprietary APIs, enabling privacy-preserving clinical decision support in real-world settings.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Medication information extraction using local large language models 96%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 95%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 95%
Similar papers in this journal
Similar papers in this journal
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 93%
Similar papers in this journal
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 94%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 93%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.