The Unreliable Judges: Assessing Reproducibility and Self-Preference Bias of LLMs as Free-Text Evaluators
Alvarez-Arenas, J. I.; mananes, D.; jimenez-carretero, d.; Sanchez-Cabo, F.
Show abstract
Large Language Models (LLMs) are transforming clinical practice and research, but their adoption requires rigorous evaluation. While human assessment is ideal, its cost has driven the widespread use of LLMs as evaluators. We introduce an open-source reciprocal framework comparing 71 human experts against six LLMs. AI evaluators show a strong self-preference bias, yet neither group reliably identified whether a response was human- or AI-generated. AI scores correlated with surface features such as length and lexical diversity, whereas human scores did not. By probing the evaluator's hidden states and applying targeted steering, we show that verbosity is a major causal driver of the bias. Moreover, shuffling question-response pairings shows that long responses keep high scores even when they no longer answer the question, whereas short ones do not, demonstrating that AI judges reward verbosity largely independently of content alignment. Finally, API-based and batch inference inflate stochasticity, underscoring the need for controlled deployment.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 94%
- DeepAction: A MATLAB toolbox for automated classification of animal behavior in video 94%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
Similar papers in this journal
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 94%
- Automated identification of abnormal infant movements from smart phone videos 92%
- Use of large language models as a scalable approach to understanding public health discourse 90%
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 93%
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 93%
- Incremental Accumulation of Linguistic Context in Artificial and Biological Neural Networks 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.