Back

Are ChatGPT and Copilot Reliable for Health Education on Statistical Testing?

Rovetta, A.; Mansournia, M. A.

2024-03-12 public and global health
10.1101/2024.03.08.24304007 medRxiv
Show abstract

The introduction of Artificial Intelligence (AI) has revolutionized daily life and scientific research, with applications ranging from writing scientific articles to clinical assistance. However, the effectiveness of AI models like ChatGPT 3.5 by Open AI and Bing Copilot GPT-4 by Microsoft in explaining complex concepts such as statistical testing is a cause for concern. This study investigates the ability of these AI models to explain fundamental statistical concepts, such as P-values, confidence intervals, and surprisals, crucial to properly inform conclusions in scientific research and public health. Our results highlight significant misconceptions in both AI models understanding and teaching of inferential statistics. These deficiencies include the mixing of incompatible statistical approaches, the nullism fallacy, the dichotomization of (statistical) significance, the incorrect interpretation of statistical measures and concepts, and an overestimation of the role of p-values and confidence intervals. Additionally, both models lack knowledge of recent alternative statistical methods like S-values and S-intervals, showing biases similar to those present in traditional statistical approaches. Given the importance of accurate statistical understanding in various sectors and the widespread integration of AI in decision-making processes, urgent intervention by OpenAI and Microsoft is necessary to update their platform databases. It is essential to align AI knowledge with the latest developments in scientific research to ensure the reliability of generated results. Collaboration with organizations such as the American Statistical Association is recommended to facilitate this process. In conclusion, this scenario underscores the need for immediate corrective action by the developing companies of such platforms. Indeed, only through continuous updates and improvements can we ensure that AI can contribute positively to scientific and technological progress.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.