Back

How easily can AI chatbots spread misinformation in audiology and otolaryngology?

Jedrzejczak, W.; Szkielkowska, A.; Raj-Koziak, D.; Wlodarczyk, E.; Skarzynski, H.; Kochanek, K.

2025-04-25 otolaryngology
10.1101/2025.04.24.25326281 medRxiv
Show abstract

BackgroundChatbots powered by large language models (LLMs) have recently emerged as prominent sources of information. However, their ability to propagate misinformation as well as information, particularly in specialized fields like audiology and otolaryngology, remains underexplored. This study aimed to evaluate the accuracy of six popular chatbots - ChatGPT, Gemini, Claude, DeepSeek, Grok, and Mistral - in response to questions framed around a range of unproven methods in audiological and otolaryngological care. MethodsA set of 50 questions was developed based on common conversations between patients and clinicians. We then posed these questions to the chatbots. We tested each chatbot 10 times to account for variable responses, producing a total of 3,000 responses. The responses were compared with correct answers based on the general opinion of 11 professionals. The consistency of the responses was evaluated by Cohens Kappa. ResultsMost chatbot responses to the majority of questions were deemed accurate. Grok consistently performed best, where its answers aligned perfectly with the opinions of the experts. Deepseek exhibited the lowest accuracy, scoring 95.8%. Mistral exhibited the lowest consistency, scoring 0.96. ConclusionsAlthough the evaluated chatbots generally avoided endorsing scientifically unsupported methods, some of the answers given could mislead and facilitate misinformation. The best performer among the group was Grok, which provided consistently accurate responses, showing it has potential for use in clinical and educational settings.

Published in OTO Open · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.