Asymmetry between warmth and clinical substance in multilingual consumer health AI
Ariel, D.; Grumberg, L. R.; Supakul, S.; Wannasri, S.; Mitchnik, I. Y.; Lev, A.; Ariyamethanon, W.; Agbarieh, M.; Miari, S.; Laban, G.; Hasid, B.
Show abstract
The same patient question can yield different clinical quality across languages. Across 504 forum-derived patient queries in six languages and four chatbots, language-matched clinicians rated responses on five clinical dimensions (1,008 ratings; 5,040 dimension scores). Patient language outweighed chatbot identity across the four clinical-substance dimensions (composite language partial {superscript 2} 0.275 vs chatbot 0.035; robust to investigator-rating exclusion: {superscript 2} 0.260) but not for empathy ({superscript 2} 0.029): clinical substance was language-associated; warmth was relatively preserved. Catastrophic safety ratings ranged 4.3-fold by language (3.6% English, 15.5% Thai and Hebrew); 62% of catastrophic ratings exceeded the English baseline (descriptive disparity). Failures were systematic and silent: none of 24 stroke responses conveyed time-criticality framing, none of 24 CO-poisoning responses challenged the familys stress framing, and 120 sentinel responses contained no confident errors. Warmth did not discriminate clinical danger (response-level empathy AUC = 0.49): consumer health AI can deliver fluent, caring tone with degraded clinical substance.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 92%
- New Model, Old Risks? Sociodemographic Bias and Adversarial Hallucinations Vulnerability in GPT-5 91%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 91%
Similar papers in this journal
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 90%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 90%
- Machine learning guided association of adverse drug reactions with in vitro target-based pharmacology 90%
Similar papers in this journal
- Zero-shot drug repurposing with geometric deep learning and clinician centered design 91%
- Actionable druggable genome-wide Mendelian randomization identifies repurposing opportunities for COVID-19 89%
- Attributes and predictors of Long-COVID: analysis of COVID cases and their symptoms collected by the Covid Symptoms Study App 89%
Similar papers in this journal
- Real-time analysis of a mass vaccination effort via an Artificial Intelligence platform confirms the safety of FDA-authorized COVID-19 vaccines 87%
- Multimodal surveillance of SARS-CoV-2 at a university enables development of a robust outbreak response framework 86%
- Application of Generative Artificial Intelligence to Utilise Unstructured Clinical Data for Acceleration of Inflammatory Bowel Disease Research 86%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.