Are ChatGPT and Copilot Reliable for Health Education on Statistical Testing?
Rovetta, A.; Mansournia, M. A.
Show abstract
The introduction of Artificial Intelligence (AI) has revolutionized daily life and scientific research, with applications ranging from writing scientific articles to clinical assistance. However, the effectiveness of AI models like ChatGPT 3.5 by Open AI and Bing Copilot GPT-4 by Microsoft in explaining complex concepts such as statistical testing is a cause for concern. This study investigates the ability of these AI models to explain fundamental statistical concepts, such as P-values, confidence intervals, and surprisals, crucial to properly inform conclusions in scientific research and public health. Our results highlight significant misconceptions in both AI models understanding and teaching of inferential statistics. These deficiencies include the mixing of incompatible statistical approaches, the nullism fallacy, the dichotomization of (statistical) significance, the incorrect interpretation of statistical measures and concepts, and an overestimation of the role of p-values and confidence intervals. Additionally, both models lack knowledge of recent alternative statistical methods like S-values and S-intervals, showing biases similar to those present in traditional statistical approaches. Given the importance of accurate statistical understanding in various sectors and the widespread integration of AI in decision-making processes, urgent intervention by OpenAI and Microsoft is necessary to update their platform databases. It is essential to align AI knowledge with the latest developments in scientific research to ensure the reliability of generated results. Collaboration with organizations such as the American Statistical Association is recommended to facilitate this process. In conclusion, this scenario underscores the need for immediate corrective action by the developing companies of such platforms. Indeed, only through continuous updates and improvements can we ensure that AI can contribute positively to scientific and technological progress.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Common misconceptions held by health researchers when interpreting linear regression assumptions, a cross-sectional study 94%
- Investigating the Role of AI Explanations in Lay Individuals’ Comprehension of Radiology Reports: A Metacognition Lense 93%
- The Mastery Rubric for Bioinformatics: supporting design and evaluation of career-spanning education and training 93%
Similar papers in this journal
- Identification, analysis and prediction of valid and false information related to vaccines from Romanian tweets 91%
- Accuracy of US CDC COVID-19 Forecasting Models 91%
- Applying machine-learning to rapidly analyse large qualitative text datasets to inform the COVID-19 pandemic response: Comparing human and machine-assisted topic analysis techniques 90%
Similar papers in this journal
- Trust in the scientific research community predicts intent to comply with COVID-19 prevention measures: An analysis of a large-scale international survey dataset 93%
- Analysis of the early Covid-19 epidemic curve in Germany by regression models with change points 91%
- Estimating the Case Fatality Ratio for COVID-19 using a Time-Shifted Distribution Analysis 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.