Mapping echocardiogram reports to a structured ontology: a task for statistical machine learning or large language models?
Subramaniam, S.; Rizvi, S.; Ramesh, R.; Sehgal, V.; Gurusamy, B.; Arif, H.; Tran, J.; Thamman, R.; Anyanwu, E.; Mastouri, R.; Mackensen, G. B.; Arnaout, R.
Show abstract
BackgroundBig data has the potential to revolutionize echocardiography by enabling novel research and rigorous, scalable quality improvement. Text reports are a critical part of such analyses, and ontology is a key strategy for promoting interoperability of heterogeneous data through consistent tagging. Currently, echocardiogram reports include both structured and free text and vary across institutions, hampering attempts to mine text for useful insights. Natural language processing (NLP) can help and includes both non-deep learning and deep-learning (e.g., large language model, or LLM) based techniques. Challenges to date in using echo text with LLMs include small corpus size, domain-specific language, and high need for accuracy and clinical meaning in model results. MethodsWe tested whether we could map echocardiography text to a structured, three-level hierarchical ontology using NLP. We used two methods: statistical machine learning (EchoMap) and one-shot inference using the Generative Pre-trained Transformer (GPT) large language model. We tested against eight datasets from 24 different institutions and compared both methods against clinician-scored ground truth. ResultsDespite all adhering to clinical guidelines, there were notable differences by institution in what information was included in data dictionaries for structured reporting. EchoMap performed best in mapping test set sentences to the ontology, with validation accuracy of 98% for the first level of the ontology, 93% for the first and second level, and 79% for the first, second, and third levels. EchoMap retained good performance across external test datasets and displayed the ability to extrapolate to examples not initially included in training. EchoMaps accuracy was comparable to one-shot GPT at the first level of the ontology and outperformed GPT at second and third levels. ConclusionsWe show that statistical machine learning can achieve good performance on text mapping tasks and may be especially useful for small, specialized text datasets. Furthermore, this work highlights the utility of a high-resolution, standardized cardiac ontology to harmonize reports across institutions.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Comparative Analysis of Privacy-Preserving Large Language Models For Automated Echocardiography Report Analysis 96%
- Natural language inference for clinical registry curation 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
Similar papers in this journal
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 96%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 93%
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 93%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 93%
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 92%
- Comparing neural language models for medical concept representation and patient trajectory prediction 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.