Diagnosis in Bytes: Comparing the Diagnostic Accuracy of Google and ChatGPT 3.5 as Diagnostic Support Tools
Guimaraes, G.; Santos Silva, C.; Contreras, J. C. Z.; Figueiredo, R. G.; Uros - Grupo de Pesquisa, ; Tiraboschi, R. B.; Gomes, C. M.; Bessa, J.
Show abstract
ObjectiveAdopting digital technologies as diagnostic support tools in medicine is unquestionable. However, the accuracy in suggesting diagnoses remains controversial and underexplored. We aimed to evaluate and compare the diagnostic accuracy of two primary and accessible internet search tools: Google and ChatGPT 3.5. MethodWe used 60 clinical cases related to urological pathologies to evaluate both platforms. These cases were divided into two groups: one with common conditions (constructed from the most frequent symptoms, following EAU and UpToDate guidelines) and another with rare disorders - based on case reports published between 2022 and 2023 in Urology Case Reports. Each case was inputted into Google Search and ChatGPT 3.5, and the results were categorized as "correct diagnosis," "likely differential diagnosis," or "incorrect diagnosis." A team of researchers evaluated the responses blindly and randomly. ResultsIn typical cases, Google achieved 53.3% accuracy, offering a likely differential diagnosis in 23.3% and errors in the rest. ChatGPT 3.5 exhibited superior performance, with 86.6% accuracy, and suggested a reasonable differential diagnosis in 13.3%, without mistakes. In rare cases, Google did not provide correct diagnoses but offered a likely differential diagnosis in 20%. ChatGPT 3.5 achieved 16.6% accuracy, with 50% differential diagnoses. ConclusionChatGPT 3.5 demonstrated higher diagnostic accuracy than Google in both contexts. The platform showed acceptable accuracy in common cases; however, limitations in rare cases remained evident.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Factors influencing plasma galectin-3 concentrations in catheter-bearing hospitalized patients 89%
- Demarcation line determination for diagnosis of gastric cancer disease range using unsupervised machine learning in magnifying narrow-band imaging 89%
- Availability and Use of Mobile Health Technology for Disease Diagnosis and Treatment Support by Health Workers in the Ashanti Region of Ghana: A Cross-sectional Survey 88%
Similar papers in this journal
- Intraoperative calculus or hemorrhage in transurethral seminal vesiculoscopy as a risk factor for recurrent hemospermia 93%
- Real-world evidence: telemedicine for complicated cases of urinary tract infection 92%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
Similar papers in this journal
- Performance of o1 pro and GPT-4 in self-assessment questions for nephrology board renewal 92%
- Emerging Applications of NLP and Large Language Models in Gastroenterology and Hepatology: A Systematic Review 91%
- Computer vision detects inflammatory arthritis in standardized smartphone photographs in an Indian patient cohort 88%
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 92%
- Artificial Intelligence Model for Analyzing Colonic Endoscopy Images to Detect Changes Associated with Irritable Bowel Syndrome 92%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.