ChatGPT Influence on Medical Decision-Making, Bias, and Equity: A Randomized Study of Clinicians Evaluating Clinical Vignettes
Goh, E.; Bunning, B.; Khoong, E.; Gallo, R.; Milstein, A.; Centola, D.; Chen, J. H.
Show abstract
In a randomized, pre-post intervention study, we evaluated the influence of a large language model (LLM) generative AI system on accuracy of physician decision-making and bias in healthcare. 50 US-licensed physicians reviewed a video clinical vignette, featuring actors representing different demographics (a White male or a Black female) with chest pain. Participants were asked to answer clinical questions around triage, risk, and treatment based on these vignettes, then asked to reconsider after receiving advice generated by ChatGPT+ (GPT4). The primary outcome was the accuracy of clinical decisions based on pre-established evidence-based guidelines. Results showed that physicians are willing to change their initial clinical impressions given AI assistance, and that this led to a significant improvement in clinical decision-making accuracy in a chest pain evaluation scenario without introducing or exacerbating existing race or gender biases. A survey of physician participants indicates that the majority expect LLM tools to play a significant role in clinical decision making.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 89%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 89%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 89%
Similar papers in this journal
- Medical history predicts phenome-wide disease onset 88%
- Deep representation learning for clustering longitudinal survival data from electronic health records 88%
- Integration of clinical characteristics, lab tests and a deep learning CT scan analysis to predict severity of hospitalized COVID-19 patients 87%
Similar papers in this journal
- Zero-shot drug repurposing with geometric deep learning and clinician centered design 89%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 87%
- Doctors and Nurses Social Media Ads Reduced Holiday Travel and COVID-19 infections: A cluster randomized controlled trial in 13 States 87%
Similar papers in this journal
- Outcomes of a Smartphone-based Application with Live Health-Coaching Post-Percutaneous Coronary Intervention 89%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 88%
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 87%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.