Using Large Language Models to Determine Reasons for Missed Colon Cancer Screening Follow-Up
Williams, C. Y. K.; Sarkar, U.; Adler-Milstein, J.; Rotenstein, L.
Show abstract
ImportanceIdentifying reasons for missed preventive care, such as follow-up colonoscopy after an abnormal stool-based colon cancer screening test, is critical for quality improvement initiatives. However, manual chart review to extract this information from unstructured clinical notes is time-consuming and costly. ObjectiveTo determine whether a large language model (LLM) can accurately extract reasons for a lack of follow-up colonoscopy after abnormal outpatient fecal immunohistochemical test (FIT) or fecal occult blood test (FOBT). DesignCross-sectional study. SettingUniversity of California, San Francisco (UCSF). ParticipantsAdult patients aged 45 years or older with an abnormal outpatient FIT/FOBT between 2012 and 2024 who did not undergo a colonoscopy within 90 days of the abnormal test. ExposureWe investigate the potential of an LLM to determine whether reasons for a lack of follow-up colonoscopy are documented in the clinical notes and whether an LLM can accurately classify those reasons into clinically meaningful categories. Main Outcomes and MeasuresAccuracy score was calculated to evaluate LLM performance against a 10% subsample manually classified by a physician reviewer. ResultsFrom a total of 2164 patients with abnormal FIT/FOBTs performed at UCSF during the study period, 355 (16.4%) underwent a colonoscopy within 90 days of the abnormal test. Among those who did not receive a colonoscopy within 90 days, 846 patients were eligible for the main analysis. Based on LLM categorization of patient note content, 270 (31.9%) patients did not have any reference to colonoscopy/colorectal cancer screening in their notes, 379 (44.8%) patients had mentions of colonoscopy/colorectal cancer screening without explicit reasons for not having a colonoscopy provided, and 197 (23.3%) patients had notes detailing explicit reasons for not having a colonoscopy. Overall LLM classification accuracy was 89.3%. The most common reasons for not having a colonoscopy included: Refused/not interested (n = 96; 35.2%), Comorbidities (n = 51; 18.7%), and Patient Unavailable (n = 46; 16.8%). Conclusions and RelevanceThis study suggests that an LLM can accurately identify and categorize reasons for the absence of follow-up colonoscopy after an abnormal FIT/FOBT. Our results suggest that LLMs have the potential to automate chart review for quality improvement initiatives.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Heterogeneity of Diagnosis and Documentation of Post-COVID Conditions in Primary Care: A Machine Learning Analysis 93%
- Clinical prediction rule for SARS-CoV-2 infection from 116 U.S. emergency departments 92%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 91%
Similar papers in this journal
- Accuracy of Medical Billing Data Against the Electronic Health Record in the Measurement of Colorectal Cancer Screening Rates 93%
- Interventions To Improve Patient Safety During The COVID-19 Pandemic: A Systematic Review 90%
- Utilisation of Remote Capillary Blood Testing in an Outpatient Clinic Setting to improve shared decision making and patient and clinician experience: a validation and pilot study 89%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 91%
- Missing data in the medical record for oncology patients: prevalence and association with outcomes 91%
- If you build it, will they use it? Use of a Digital Assistant for Self-Reporting of COVID-19 Rapid Antigen Test Results during Large Nationwide Community Testing Initiative 91%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 92%
- Observer: Creation of a Novel Multimodal Dataset for Outpatient Care Research 92%
Similar papers in this journal
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 93%
- Cancer risk algorithms in primary care: can they improve risk estimates and referral decisions? 93%
- The prevalence of SARS-CoV-2 infection and other public health outcomes during the BA.2/BA.2.12.1 surge, New York City, April-May 2022 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.