Back

Multi-LLM Disagreement as a Scalable Detector of Human Annotation Errors in Structured Data from Clinical Free-Text

Wittlinger, S.; Meerjansen, J.; Wolf, F.; Wiest, I. C.; Ebert, M. P.; Siegel, F.; Belle, S.

2026-05-06 health systems and quality improvement
10.64898/2026.05.04.26352392 medRxiv
Show abstract

ObjectiveStructured extraction from clinical free-text depends on human annotators whose labels are susceptible to errors and knowledge-driven mistakes; exhaustive quality control is impractical at scale. We evaluate whether disagreement among multiple locally hosted large language models (LLMs) can prioritize human annotations for targeted review. MethodsMultiple LLMs independently extract the same set of structured variables annotated by a human reviewer. For each annotation, an agreement score counts the LLMs matching the human label. Using four locally hosted LLMs (Gemma 3 27B, DeepSeek-R1 70B, GPT-OSS 120B, Mistral Large 3), we evaluated this approach on 910 German-language colonoscopy reports describing endoscopic mucosal resection, with five structured variables per case (anatomical location, two diameters, resection technique, multiple polyps), yielding 4,550 annotations and a 377-case adjudication sample. A stratified sample oversampling low-agreement strata was adjudicated blinded by an experienced reviewer and analyzed with prevalence-adjusted estimates ResultsHuman error rates rose as LLM agreement fell, from 0% at scores 3-4 to 76% at score 0. The lowest-agreement stratum was only 6.5% of annotations yet concentrated an estimated 80% of errors. The multi-LLM disagreement score achieved a prevalence-adjusted AUC-ROC of 0.991 (95% CI 0.987-0.994) and AUC-PR of 0.893 (95% CI 0.851-0.929) for error detection. DiscussionMulti-LLM disagreement outperformed single models and provided graded operating points for risk-stratified review. ConclusionMulti-LLM disagreement provides a scalable quality-control signal for targeted review of the highest-yield cases. Because all models run locally, the framework is GDPR-compliant; its language- and task-agnostic design supports application across clinical domains.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 18%
9.8%
2
PLOS ONE
5266 papers in training set
Top 19%
9.7%
3
Communications Medicine
113 papers in training set
Top 0.2%
7.9%
4
npj Digital Medicine
118 papers in training set
Top 0.8%
7.3%
5
Clinical Trials
11 papers in training set
Top 0.1%
6.7%
6
Biometrics
23 papers in training set
Top 0.1%
4.3%
7
Scientific Reports
3612 papers in training set
Top 26%
4.0%
8
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.6%
3.2%
50% of probability mass above
9
Journal of the American Medical Informatics Association
71 papers in training set
Top 1%
2.4%
10
Nature
645 papers in training set
Top 6%
2.1%
11
BMC Medical Research Methodology
47 papers in training set
Top 0.6%
2.1%
12
Nature Medicine
125 papers in training set
Top 1%
1.9%
13
BMJ Open Quality
17 papers in training set
Top 0.3%
1.9%
14
The Lancet Digital Health
25 papers in training set
Top 0.3%
1.7%
15
Healthcare
17 papers in training set
Top 0.4%
1.5%
16
Journal of Medical Internet Research
87 papers in training set
Top 2%
1.5%
17
eBioMedicine
183 papers in training set
Top 3%
1.3%
18
Journal of Biomedical Informatics
47 papers in training set
Top 1.0%
1.1%
19
Journal of Clinical Epidemiology
31 papers in training set
Top 0.6%
1.1%
20
Med
39 papers in training set
Top 0.4%
1.1%
21
JAMA Network Open
130 papers in training set
Top 3%
1.1%
22
JMIR Medical Informatics
18 papers in training set
Top 0.6%
1.1%
23
PLOS Digital Health
106 papers in training set
Top 3%
1.1%
24
JAMIA Open
42 papers in training set
Top 1%
1.1%
25
iScience
1154 papers in training set
Top 25%
1.1%
26
BMJ Health & Care Informatics
15 papers in training set
Top 0.8%
1.1%
27
Annals of Internal Medicine
28 papers in training set
Top 0.5%
1.0%
28
Frontiers in Medicine
120 papers in training set
Top 3%
1.0%
29
Clinical Pharmacology & Therapeutics
25 papers in training set
Top 0.3%
1.0%
30
Modern Pathology
22 papers in training set
Top 0.3%
1.0%