Implementing a Resource-Light and Low-Code Large Language Model System for Information Extraction from Mammography Reports: A Case Study
Dennstaedt, F.; Fauser, S.; Cihoric, N.; Schmerder, M.; Lombardo, P.; Cereghetti De Marchi, G.; von Daeniken, S.; Minder, T.; Meyer, J.; Chiang, L.; Gaio, R.; Lerch, L.; Filchenko, I.; Reichenpfader, D.; Denecke, K.; Vojvodic, C.; Tatalovic, I.; Sander, A.; Hastings, J.; Aebersold, D. M.; von Tengg, H.; Nairz, K.
Show abstract
BackgroundLarge Language Models (LLMs) have been successfully used to extract structured data from free-text radiology reports. Most of current studies were conducted with private models accessed via Application Programming Interface (API). We aimed to evaluate the feasibility of using open-source LLMs, deployed on limited local hardware resources for extraction of structured information from free-text mammography reports, according to a Common Data Elements (CDE)-based framework. MethodsSeventy-nine CDEs were defined by an interdisciplinary expert panel, reflecting real-world reporting practice. Sixty-one reports were classified by two independent researchers with 1533 classifications assigned to establish ground truth. Five different open-source LLMs deployable on a single GPU were used for data extraction using the general-classifier Python package. Extractions were performed for two different prompt approaches with classification metrics calculated overall and on subgroups. Additional analyses were conducted using thresholds for the relative probability of classifications. ResultsHigh inter-rater agreement was observed between manual classifiers (Cohens Kappa 0.83). Using default prompts, the LLMs achieved accuracies of 59.23-72.86%. Adapting prompts to better explain classification tasks improved performance for all models, with accuracies of 64.71-85.32%. Setting certainty thresholds further improved accuracies to >90% but reduced the coverage rate to <50%. ConclusionLocally deployed open-source LLMs can effectively extract information from mammography reports with good accuracy, addressing data privacy concerns while maintaining compatibility with limited computational resources. Prompt engineering substantially increases performance, highlighting the importance of optimization in clinical applications. Using a CDE-based framework provides clear semantics and structure, facilitating interoperability and consistent data extraction.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Health Data Nexus: An Open Data Platform for AI Research and Education in Medicine 95%
- New implementation of data standards for AI research in precision oncology. Experience from EuCanImage 94%
- Strategies and Techniques for Quality Control and Semantic Enrichment with Multimodal Data: A Case Study in Colorectal Cancer with eHDPrep 93%
Similar papers in this journal
- Implementation and prospective real-time evaluation of a generalized system for in-clinic deployment and validation of machine learning models in radiology 95%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 95%
- Classification of Hyper-scale Multimodal Imaging Datasets 94%
Similar papers in this journal
Similar papers in this journal
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 94%
- On evaluation metrics for medical applications of artificial intelligence 94%
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 92%
Similar papers in this journal
- Model uncertainty estimates for deep learning mammographic density prediction using ordinal and classification approaches 91%
- SimBSI: An open-source Simulink library for developingclosed-loop brain signal interfaces in animals and humans 90%
- Breast density prediction from low and standard dose mammograms using deep learning: effect of image resolution and model training approach on prediction quality 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.