How easily can AI chatbots spread misinformation in audiology and otolaryngology?
Jedrzejczak, W.; Szkielkowska, A.; Raj-Koziak, D.; Wlodarczyk, E.; Skarzynski, H.; Kochanek, K.
Show abstract
BackgroundChatbots powered by large language models (LLMs) have recently emerged as prominent sources of information. However, their ability to propagate misinformation as well as information, particularly in specialized fields like audiology and otolaryngology, remains underexplored. This study aimed to evaluate the accuracy of six popular chatbots - ChatGPT, Gemini, Claude, DeepSeek, Grok, and Mistral - in response to questions framed around a range of unproven methods in audiological and otolaryngological care. MethodsA set of 50 questions was developed based on common conversations between patients and clinicians. We then posed these questions to the chatbots. We tested each chatbot 10 times to account for variable responses, producing a total of 3,000 responses. The responses were compared with correct answers based on the general opinion of 11 professionals. The consistency of the responses was evaluated by Cohens Kappa. ResultsMost chatbot responses to the majority of questions were deemed accurate. Grok consistently performed best, where its answers aligned perfectly with the opinions of the experts. Deepseek exhibited the lowest accuracy, scoring 95.8%. Mistral exhibited the lowest consistency, scoring 0.96. ConclusionsAlthough the evaluated chatbots generally avoided endorsing scientifically unsupported methods, some of the answers given could mislead and facilitate misinformation. The best performer among the group was Grok, which provided consistently accurate responses, showing it has potential for use in clinical and educational settings.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- What do blind people "see" with retinal prostheses? Observations and qualitative reports of epiretinal implant users 92%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 92%
Similar papers in this journal
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 92%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 91%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 91%
Similar papers in this journal
- Performance of ChatGPT in pediatric audiology as rated by students and experts 97%
- Artificial Intelligence in laryngeal endoscopy: Systematic Review and Meta-Analysis 91%
- Aerodigestoscopy (ADS): A retrospective examination of the feasibility, safety, and comfort of a new procedure for the evaluation of physiological disorders of the aerodigestive tract 90%
Similar papers in this journal
- Cross-modal sensory boosting to improve high-frequency hearing loss 94%
- Challenges in Implementing a Mobile AI Chatbot Intervention for Depression Among Youth on Psychiatric Waiting Lists: A Randomized Control Study Termination Report 89%
- Learning from the resilience of hospitals and their staff to the COVID-19 pandemic: a scoping review 89%
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.