Back

Automated Identification of Complex Percutaneous Coronary Intervention from Cardiac Catheterization Reports Using Large Language Models

Bhatt, N.; Warner, F.; Miao, J.; Thakker, R.; Joodi, G.; Cantero-Schaffer, P.; Huang, C.; Krumholz, H.; Murugiah, K.

2026-08-10 cardiovascular medicine
10.64898/2026.08.05.26359802 medRxiv
Show abstract

Background: Manual abstraction of complex percutaneous coronary intervention (PCI) variables from cardiac catheterization reports is labor-intensive and limits scalable cardiovascular research. Large language models (LLMs) may enable automated extraction of procedural data, but their performance remains uncertain. Methods: We evaluated three open-source LLMs (Llama 3.3 70B, Meditron-7B, and BioMistral-7B) using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health system. Models were tasked to identify if a procedure note was a PCI procedure, and extract variables used to classify PCI as complex using predefined criteria, including 3 vessels treated, [≥]3 treated lesions, bifurcation PCI with two stents, chronic total occlusion, [≥]3 stents, and total stent length [≥]60 mm. Results: The evaluation cohort included 1,412 clinical notes of which 596 were PCI procedures. Llama 3.3 70B consistently outperformed both domain-specific models across nearly all extraction tasks. For PCI identification, Llama 3 70B had 100.0% sensitivity, 93.8% specificity, 92.1% positive predictive value, 100.0% negative predictive value, 96.4% accuracy, and an F1 score of 95.9%. For complex PCI classification, among 590 evaluable PCI reports, sensitivity was 97.7%, specificity was 80.1%, positive predictive value was 57.6%, negative predictive value was 99.2%, accuracy was 83.9%, and the F1 score was 72.5%. Variables that were explicitly documented, including stent number, stent length, and adjunctive device use, were extracted with high accuracy, whereas performance was lower for variables requiring contextual reasoning, including lesion counting, bifurcation PCI, and chronic total occlusion. Conclusion: High-capacity open-source LLMs can accurately extract complex PCI variables from free-text catheterization reports, supporting LLM-enabled automated phenotyping to reduce manual abstraction and facilitate scalable cardiovascular research.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.