Large Language Models Readability Classification: A Variability Analysis of Sources and Metrics
Corrale de Matos, H. G.; Wasmann, J.-W. A.; Catalani Morata, T.; de Freitas Alvarenga, K.; Bornia Jacob, L. C.
Show abstract
AbstractAccurate health information is ineffective if patients cannot understand it. Large Language Model (LLM) health research values veridical precision; however, linguistic accessibility remains an under-examined component of output quality and usability. This study investigated two sources of variability in readability classification: differences across LLM systems and across readability metrics. The analysis tested 1,120 data points from seven systems in English and Portuguese, comparing baseline responses with a Wikipedia-grounded condition. Content was assessed using five standard readability metrics that measure distinct aspects of text complexity. Systems were statistically homogeneous at baseline but became significantly heterogeneous under Wikipedia grounding, indicating variability in the combination of Retrieval-Augmented Generation (differential readability effects of the same source-grounding instruction across systems). Significant metric variability was observed in all conditions, showing that readability metrics are not interchangeable. Although retrieval grounding is commonly used to improve accuracy, our findings show a trade-off: verified-source grounding can yield inconsistent readability. Therefore, evaluation protocols should use transparent, vendor-agnostic criteria, with metric-specific and language-aware thresholds, and be applied whenever models or grounding configurations change to support accessible cross-language health communication.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Evaluation of Large Language Models in Medical Examinations:A Scoping Review Protocol 93%
- Mapping the quality of Norwegian health information - Does it facilitate informed choices? 92%
- Investigating the Role of AI Explanations in Lay Individuals’ Comprehension of Radiology Reports: A Metacognition Lense 92%
Similar papers in this journal
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 93%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 92%
- User Experience Evaluation of Cogscreen for Screening Mild Cognitive Impairment: Formative and Summative Evaluation 91%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 93%
- The development of a World Health Organization transdiagnostic chatbot intervention for distressed adolescents and young adults 91%
- Implementing Home-Based Digital Health in Rural Canada: A Scoping Review 90%
Similar papers in this journal
- Has the pandemic enhanced and sustained digital health-seeking behaviour? A big data interrupted time-series analysis of Google Trends 91%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 91%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.