Customizing GPT-4 for clinical information retrieval from standard operating procedures
Muti, H. S.; Loeffler, C. M. L.; Lessmann, M. E.; Stueker, E. H.; Kirchberg, J.; von Bonin, M.; Kolditz, M.; Ferber, D.; Egger-Heidrich, K.; Merboth, F.; Stange, D.; Distler, M.; Kather, J. N.
Show abstract
BackgroundThe increasing complexity of medical knowledge necessitates efficient and reliable information access systems in clinical settings. For quality purposes, most hospitals use standard operating procedures (SOPs) for information management and implementation of local treatment standards. However, in clinical routine, this information is not always easily accessible. Customized Large Language Models (LLMs) may offer a tailored solution, but need thorough evaluation prior to clinical implementation. ObjectiveTo customize an LLM to retrieve information from hospital-specific SOPs, to evaluate its accuracy for clinical use and to compare different prompting strategies and large language models. MethodsWe customized GPT-4 with a predefined system prompt and 10 SOPs from four departments at the University Hospital Dresden. The models performance was evaluated through 30 predefined clinical questions of varying degree of detail, which were assessed by five observers with different levels of medical expertise through simple and interactive question-and-answering (Q&A). We assessed answer completeness, correctness and sufficiency for clinical use and the impact of prompt design on model performance. Finally, we compared the performance of GPT-4 with Claude-3-opus. ResultsInteractive Q&A yielded the highest rate of completeness (80%), correctness (83%) and sufficiency (60%). Acceptance of the LLMs answer was higher among early-career medical staff. Degree of detail of the question prompt influenced answer accuracy, with intermediate-detail prompts achieving the highest sufficiency rates. Comparing LLMs, Claude-3-opus outperformed GPT-4 in providing sufficient answers (70.0% vs. 36.7%) and required fewer iterations for satisfactory responses. Both models adhered to the system prompt more effectively in the self-coded pipeline than in the browser application. All observers showed discrepancies between correctness and accuracy of the answers, which rooted in the representation of information in the SOPs. ConclusionInteractively querying customized LLMs can enhance clinical information retrieval, though expert oversight remains essential to ensure a safe application of this technology. After broader evaluation and with basic knowledge in prompt engineering, customized LLMs can be an efficient, clinically applicable tool.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 96%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 95%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 95%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 95%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 97%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 96%
- Assessment of Accuracy and Safety of LabTest Checker (LTC-AI) 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.