Examining ChatGPTs Ability to Answer Cirrhosis Related Questions in Arabic Compared to English
Samaan, J. S.; Yeo, Y. H.; Ng, W. H.; Ting, P.-S.; Trivedi, H.; Vipani, A.; Yang, J. D.; Liran, O.; Spiegel, B.; Kuo, A.; Ayoub, W.
Show abstract
Background and Study AimsCirrhosis is a chronic progressive disease which requires complex care. Its incidence is rising in the Arab countries making it the 7th leading cause of death in the Arab League in 2010. ChatGPT is a large language model with a growing body of literature demonstrating its ability to answer clinical questions. We examined ChatGPTs accuracy in responding to cirrhosis related questions in Arabic and compared its performance to English. Materials and MethodsChatGPTs responses to 91 questions in Arabic and English were graded by a transplant hepatologist fluent in both languages. Accuracy of responses was assessed using the scale: 1. Comprehensive, 2. Correct but inadequate, 3. Mixed with correct and incorrect/outdated data, and 4. Completely incorrect. Accuracy of Arabic compared to English responses was assessed using the scale: 1. Arabic response is more accurate, 2. Similar accuracy, 3. Arabic response is less accurate. ResultsThe model provided 22 (24.2%) comprehensive, 44 (48.4%) correct but inadequate, 13 (14.3%) mixed with correct and incorrect/outdated data and 12 (13.2%) completely incorrect Arabic responses. When comparing the accuracy of Arabic and English responses, 9 (9.9%) of the Arabic responses were graded as more accurate, 52 (57.1%) similar in accuracy and 30 (33.0%) as less accurate compared to English. ConclusionChatGPT has the potential to serve as an adjunct source of information for Arabic speaking patients with cirrhosis. The model provided correct responses in Arabic to 72.5% of questions, although its performance in Arabic was less accurate than in English. The model produced completely incorrect responses to 13.2% of questions, reinforcing its potential role as an adjunct and not replacement of care by licensed healthcare professionals. Future studies to refine this technology are needed to help Arabic speaking patients with cirrhosis across the globe understand their disease and improve their outcomes.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 92%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 91%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 91%
Similar papers in this journal
- Perceptions of Complementary, Alternative, and Integrative Medicine: Insights from a Large-Scale International Cross-Sectional Survey of Surgery Researchers and Clinicians 91%
- Improvement of Survival Outcomes of Cholangiocarcinoma by Ultrasonography Surveillance: Multicenter Retrospective Cohorts 89%
- Self-learning on COVID-19 among medical students and their preparedness to participate in government's COVID-19 response in Bhutan: a cross-sectional study 89%
Similar papers in this journal
- Impact of Losartan on Portal hypertension and Liver Cirrhosis: A Systematic Review 90%
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 90%
- “One more tool in the tool belt”: A qualitative interview study investigating patient and clinician opinions on the integration of psychometrics into routine testing for disorders of gut-brain interaction 90%
Similar papers in this journal
- A cross-sectional audit and survey of Open Science and Data Sharing practices at The Montreal Neurological Institute-Hospital 89%
- Insights into the Datasets, Tools, and Training Needs of the AnVIL Community: 2024 88%
- Glibenclamide, ATP and Metformin Increases the Expression of Human Bile Salt Export Pump ABCB11 88%
Similar papers in this journal
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 92%
- Application of physiological network mapping in the prediction of survival in critically ill patients with acute liver failure 90%
- Researcher and Clinician Preferences for a Journal Transparency Tool: A Mixed-Methods Survey and Focus Group Study 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.