Back

Detecting Cardiac Amyloidosis in Italian Cardiology Reports: Structured Variable Extraction versus Direct Free-Text Analysis

Mazzucato, S.; Sartiano, D.; Vergaro, G.; Dalmiani, S.; Emdin, M.; Micera, S.; Oddo, C. M.; Passino, C.; Moccia, S.; Bandini, A.; Seinen, T.

2026-01-23 health informatics
10.64898/2026.01.22.26344604 medRxiv
Show abstract

BackgroundEarly and accurate identification of cardiac amyloidosis improves patient outcomes, yet relevant evidence is frequently hidden in free-text records. This study assesses whether structured variable extraction or direct free-text analysis more reliably identifies patients with cardiac amyloidosis, with the goal of informing clinical decision support strategies. MethodsWe extracted 21 clinical variables from 432 Italian patient records using supervised and prompt-based methods with both proprietary and locally-deployable computational models. Classification performance was evaluated by comparing extracted data with gold-standard manual annotations. Two feature sets were tested: general clinical variables and amyloidosis-specific risk factors. Additionally, we evaluated direct zero-shot prediction on unstructured clinical notes. ResultsFor entity extraction, GPT-4.1-mini achieved F1=0.96, comparable to supervised SpaCy (F1=0.95) and GPT-4o (F1=0.94). Local open-source models like Qwen2.5 reached F1=0.94. For cardiac amyloidosis classification, machine learning models using full extracted features matched gold-standard results (SauerkrautLM-Gemma: F1=0.80 vs. gold: F1=0.82). General clinical features alone yielded lower performance (F1=0.68), highlighting that amyloidosis-specific risk factors in unstructured text provide discriminative diagnostic value. Zero-shot direct predictions outperformed supervised feature-based approaches (MedGemma: F1=0.92). ConclusionsAutomated extraction and zero-shot prediction effectively structure Italian EHRs and identify amyloidosis patients without manual annotation. Domain-specific risk factors in free-text notes provide substantial predictive value. Italian hospitals can potentially deploy locally-deployable models to screen cardiac amyloidosis without manual annotation or proprietary APIs, enabling privacy-preserving clinical decision support in real-world settings.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.