Can LLMs Improve Healthcare Delivery? Evidence from Physician Review and Objective Testing
Abaluck, J.; Pless, R.; Ravi, N.; Sautmann, A.; Schwartz, A.
Show abstract
We deployed large language model (LLM) decision support for health workers at two outpatient clinics in Nigeria. For each patient, health workers drafted care plans that were optionally revised after LLM feedback. We compared unassisted and assisted plans using blinded randomized assessments by on-site physicians who evaluated and treated the same patients and using results from laboratory tests for common conditions. Academic physicians performed blinded retrospective reviews of a subset of notes. In response to LLM feedback, health workers changed their prescribing for more than half of patients. Health workers reported high satisfaction with LLM feedback and retrospective academic reviewers rated LLM-assisted plans more favorably. However, on-site physicians observed little to no improvement in diagnostic alignment or treatment decisions. Laboratory testing showed mixed effects of LLM-assistance, which removed negative tests for malaria but added them for urinary tract infection and anemia, with no significant increase in the detection rates for the tested conditions. This highlights a gap between chart-based reviews and real-world clinical relevance that may be especially important in evaluating the effectiveness of LLM-based interventions.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Spatio-temporal modelling of referrals to outpatient respiratory clinics in the integrated care system of the Morecambe Bay area, England 93%
- The SENTINEL study of differentiated service delivery models for HIV treatment in Malawi, South Africa, and Zambia: research protocol for a prospective cohort study 92%
- ‘Leading from the front’ implementation strategies increase the success of influenza vaccination drives among healthcare workers: A reanalysis of Systematic Review evidence using Intervention Component Analysis (ICA) and Qualitative Comparative Analysis (QCA) 92%
Similar papers in this journal
- Essential Indicators of Quality in Primary Care Settings: An Evidence-Based, Structured, Expert Approach 93%
- Cost savings in male circumcision post-operative care continuum in rural and urban South Africa: Evidence on the importance of initial counselling and daily SMS 92%
- Title-Impact of combined medication payment management policies on population health performance 92%
Similar papers in this journal
- Measure what matters: counts of hospitalized patients are a better metric for health system capacity planning for a reopening 94%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 91%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 91%
Similar papers in this journal
- Unequal Recovery in Colorectal Cancer Screening Following the COVID-19 Pandemic: A Comparative Microsimulation Analysis 91%
- A modular approach to integrating multiple data sources into real-time clinical prediction for pediatric diarrhea 91%
- COVID-19 clusters in schools: frequency, size, and transmission rates from crowdsourced exposure reports 91%
Similar papers in this journal
- Bias reduction and inference for electronic health record data under selection and phenotype misclassification: three case studies 93%
- Sensitivity to missing not at random dropout in clinical trials: use and interpretation of the Trimmed Means Estimator 93%
- Toward Evaluation of Disseminated Effects of Medications for Opioid Use Disorder within Provider-Based Clusters Using Routinely-Collected Health Data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.