Development of a privacy preserving large language model for automated data extraction from thyroid cancer pathology reports
Lee, D.; Vaid, A.; Menon, K.; Freeman, R.; Matteson, D.; Marin, M.; Nadkarni, G.
Show abstract
BackgroundPopularized by ChatGPT, large language models (LLM) are poised to transform the scalability of clinical natural language processing (NLP) downstream tasks such as medical question answering (MQA) and may enhance the ability to rapidly and accurately extract key information from clinical narrative reports. However, the use of LLMs in the healthcare setting is limited by cost, computing power and concern for patient privacy. In this study we evaluate the extraction performance of a privacy preserving LLM for automated MQA from surgical pathology reports. Methods84 thyroid cancer surgical pathology reports were assessed by two independent reviewers and the open-source FastChat-T5 3B-parameter LLM using institutional computing resources. Longer text reports were converted to embeddings. 12 medical questions for staging and recurrence risk data extraction were formulated and answered for each report. Time to respond and concordance of answers were evaluated. ResultsOut of a total of 1008 questions answered, reviewers 1 and 2 had an average concordance rate of responses of 99.1% (SD: 1.0%). The LLM was concordant with reviewers 1 and 2 at an overall average rate of 88.86% (SD: 7.02%) and 89.56% (SD: 7.20%). The overall time to review and answer questions for all reports was 206.9, 124.04 and 19.56 minutes for Reviewers 1, 2 and LLM, respectively. ConclusionA privacy preserving LLM may be used for MQA with considerable time-saving and an acceptable accuracy in responses. Prompt engineering and fine tuning may further augment automated data extraction from clinical narratives for the provision of real-time, essential clinical insights.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 91%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 91%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 91%
Similar papers in this journal
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 88%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 88%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 88%
Similar papers in this journal
- Weakly supervised learning for multi-organ adenocarcinoma classification in whole slide images 92%
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 90%
- Classification performance bias between training and test sets in a limited mammography dataset 90%
Similar papers in this journal
- Scientific hypothesis generation process in clinical research: a secondary data analytic tool versus experience study protocol 89%
- A conversational agent for providing personalized PrEP support: Protocol for chatbot implementation 87%
- Ensuring a successful transition from Pap to HPV-based primary screening in Canada: a study protocol to investigate the psychosocial correlates of women’s screening intentions 86%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.