The Illusion of Understanding: A Randomized Controlled Trial of LLM-Generated Lay Summaries of Brain MRI Reports
Le Guellec, B.; Bentegeac, R.; Tran, V.-T.; El Homsi, M.; Amouyel, P.; Kuchcinski, G.; Hamroun, A.
Show abstract
Background: Large language models have been proposed to improve patient comprehension of radiology reports. However, whether they improve objective understanding remains unproven. Purpose: To evaluate the effect of appending an LLM-generated lay summary to brain MRI reports on objective and subjective patient comprehension in a randomized controlled trial. Materials and Methods: In this randomized controlled trial, 2,727 adult participants from the ComPaRe e-cohort were randomly assigned to interpret six standardized brain MRI reports for headache, presented either in their native format (control; n = 1,401) or appended with a lay summary generated by an open-weights LLM (Mistral Small 3.2) (intervention; n = 1,326). The primary outcome was objective comprehension, defined as the rate of correct classification of whether the report provided a probable explanation for the headache, with ground truth established by four-radiologist consensus. Secondary outcomes included satisfaction, subjective comprehension, anxiety, and willingness to contact a healthcare professional. Generalized estimating equations accounted for repeated within-participant observations. Results: A total of 2,727 participants (mean age, 52 years +/- 15; 75.2% women) were evaluated. Objective comprehension did not differ between arms (58.3% vs 59.4%; odds ratio (OR) 0.97; 95% CI: 0.90-1.06; P = .54). The intervention significantly improved overall satisfaction (64.9% vs 36.7%; OR 3.26; 95% CI: 2.93-3.64; P < .001) and subjective comprehension (50.3% vs 24.0%; OR 3.17; 95% CI: 2.82-3.56; P < .001). High anxiety was modestly reduced (25.1% vs 26.6%; OR 0.92; P = .037). The effect on objective comprehension varied by report type (P for interaction < .001): summaries improved comprehension of symptom-explaining reports (42.4% vs 37.4%; P < .001) but reduced it for normal reports (72.5% vs 76.6%; P = .001). Conclusion: LLM-generated lay summaries appended to brain MRI reports improved patient satisfaction and subjective comprehension but did not improve objective comprehension, indicating a gap between perceived and actual understanding that should be addressed before clinical integration.
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Digital Health Tools for the Passive Monitoring of Depression: A Systematic Review of Methods 90%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 90%
- Digital Health Interventions in Palliative Care: A Systematic Meta-Review and Evidence Synthesis 90%
Similar papers in this journal
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 90%
- Pairwise learning of MRI scans using a convolutional Siamese network for prediction of knee pain 87%
Similar papers in this journal
Similar papers in this journal
- Clinical practice guidelines and recommendations in the context of the COVID-19 pandemic: systematic review and critical appraisal 91%
- Use and appropriateness of Reporting Guidelines in physical therapy research: a protocol for a meta-research study 91%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.