GPT-4 outperforms ChatGPT in answering non-English questions related to cirrhosis
Yeo, Y. H.; Samaan, J. S.; Ng, W. H.; Ma, X.; Ting, P.-S.; Kwak, M.-S.; Panduro, A.; Lizaola-Mayo, B.; Trivedi, H.; Vipani, A.; Ayoub, W.; Yang, J. D.; Liran, O.; Spiegel, B.; Kuo, A.
Show abstract
Background and ObjectivesArtificial intelligence is increasingly being employed in healthcare, raising concerns about the exacerbation of disparities. This study evaluates ChatGPT and GPT-4s ability to comprehend and respond to cirrhosis-related questions in English, Korean, Mandarin, and Spanish, addressing language barriers that may impact patient care. MethodsA set of 36 cirrhosis-related questions were translated into Korean, Mandarin, and Spanish and prompted to both ChatGPT and GPT-4 models. Non-English responses were graded by native-speaking hepatologists on accuracy and similarity to English responses. Chi-square tests were used to compare the proportions of grading between ChatGPT and GPT-4. ResultsGPT-4 showed a marked improvement in the proportion of comprehensive and correct answers compared to ChatGPT across all four languages (p<0.05). GPT-4 demonstrated enhanced accuracy and avoided erroneous responses evident in ChatGPTs output. Significant improvement was observed in Mandarin and Korean subgroups, with a smaller quality gap between English and non-English responses in GPT-4 compared to ChatGPT. ConclusionsGPT-4 exhibited significantly higher accuracy in English and non-English cirrhosis-related questions, highlighting its potential for more accurate and reliable language model applications in diverse linguistic contexts. These advancements have important implications for patients with language discordance, contributing to equalizing health literacy on a global scale.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 95%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
Similar papers in this journal
Similar papers in this journal
- Impact of Losartan on Portal hypertension and Liver Cirrhosis: A Systematic Review 90%
- Benchmarking Deep Learning-based Image Retrieval of Oral Tumor Histology 90%
- The Impact of Fasting the Holy Month of Ramadan on Colorectal Cancer Patients and Two Tumor Biomarkers: A Tertiary-Care Hospital Experience 90%
Similar papers in this journal
- Improvement of Survival Outcomes of Cholangiocarcinoma by Ultrasonography Surveillance: Multicenter Retrospective Cohorts 91%
- Perceptions of Complementary, Alternative, and Integrative Medicine: Insights from a Large-Scale International Cross-Sectional Survey of Surgery Researchers and Clinicians 91%
- ChatGPT achieves comparable accuracy to specialist physicians in predicting the efficacy of high-flow oxygen therapy 90%
Similar papers in this journal
- Using Automated-Machine Learning to Predict COVID-19 Patient Survival: Identify Influential Biomarkers 92%
- Clinical Characteristics And Prognostic Factors For ICU Admission Of Patients With COVID-19 Using Machine Learning And Natural Language Processing 92%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.