The Diagnostic and Triage Accuracy of the GPT-3 Artificial Intelligence Model
Levine, D. M.; Tuwani, R.; Kompa, B.; Varma, A.; Finlayson, S. G.; Mehrotra, A.; Beam, A.
Show abstract
ImportanceArtificial intelligence (AI) applications in health care have been effective in many areas of medicine, but they are often trained for a single task using labeled data, making deployment and generalizability challenging. Whether a general-purpose AI language model can perform diagnosis and triage is unknown. ObjectiveCompare the general-purpose Generative Pre-trained Transformer 3 (GPT-3) AI models diagnostic and triage performance to attending physicians and lay adults who use the Internet. DesignWe compared the accuracy of GPT-3s diagnostic and triage ability for 48 validated case vignettes of both common (e.g., viral illness) and severe (e.g., heart attack) conditions to lay people and practicing physicians. Finally, we examined how well calibrated GPT-3s confidence was for diagnosis and triage. Setting and ParticipantsThe GPT-3 model, a nationally representative sample of lay people, and practicing physicians. ExposureValidated case vignettes (<60 words; <6th grade reading level). Main Outcomes and MeasuresCorrect diagnosis, correct triage. ResultsAmong all cases, GPT-3 replied with the correct diagnosis in its top 3 for 88% (95% CI, 75% to 94%) of cases, compared to 54% (95% CI, 53% to 55%) for lay individuals (p<0.001) and 96% (95% CI, 94% to 97%) for physicians (p=0.0354). GPT-3 triaged (71% correct; 95% CI, 57% to 82%) similarly to lay individuals (74%; 95% CI, 73% to 75%; p=0.73); both were significantly worse than physicians (91%; 95% CI, 89% to 93%; p<0.001). As measured by the Brier score, GPT-3 confidence in its top prediction was reasonably well-calibrated for diagnosis (Brier score = 0.18) and triage (Brier score = 0.22). Conclusions and RelevanceA general-purpose AI language model without any content-specific training could perform diagnosis at levels close to, but below physicians and better than lay individuals. The model was performed less well on triage, where its performance was closer to that of lay individuals.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 94%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 92%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.