Back

GPT-4 outperforms ChatGPT in answering non-English questions related to cirrhosis

Yeo, Y. H.; Samaan, J. S.; Ng, W. H.; Ma, X.; Ting, P.-S.; Kwak, M.-S.; Panduro, A.; Lizaola-Mayo, B.; Trivedi, H.; Vipani, A.; Ayoub, W.; Yang, J. D.; Liran, O.; Spiegel, B.; Kuo, A.

2023-05-05 gastroenterology
10.1101/2023.05.04.23289482 medRxiv
Show abstract

Background and ObjectivesArtificial intelligence is increasingly being employed in healthcare, raising concerns about the exacerbation of disparities. This study evaluates ChatGPT and GPT-4s ability to comprehend and respond to cirrhosis-related questions in English, Korean, Mandarin, and Spanish, addressing language barriers that may impact patient care. MethodsA set of 36 cirrhosis-related questions were translated into Korean, Mandarin, and Spanish and prompted to both ChatGPT and GPT-4 models. Non-English responses were graded by native-speaking hepatologists on accuracy and similarity to English responses. Chi-square tests were used to compare the proportions of grading between ChatGPT and GPT-4. ResultsGPT-4 showed a marked improvement in the proportion of comprehensive and correct answers compared to ChatGPT across all four languages (p<0.05). GPT-4 demonstrated enhanced accuracy and avoided erroneous responses evident in ChatGPTs output. Significant improvement was observed in Mandarin and Korean subgroups, with a smaller quality gap between English and non-English responses in GPT-4 compared to ChatGPT. ConclusionsGPT-4 exhibited significantly higher accuracy in English and non-English cirrhosis-related questions, highlighting its potential for more accurate and reliable language model applications in diverse linguistic contexts. These advancements have important implications for patients with language discordance, contributing to equalizing health literacy on a global scale.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.