Back

Detection without calibration: benchmarking domestic and international large language models for quality control of Mandarin 18F-FDG PET/CT reports

Wang, J.; Tang, W.; Ma, X.; Yan, H. m.; Yuan, Y.

2026-06-26 radiology and imaging
10.64898/2026.06.24.26356406 medRxiv
Show abstract

Large language models (LLMs) are increasingly used for automated quality control (QC) of radiology reports. However, the reliability of LLMs on reports in Mandarin, and the relative performance of domestic versus international flagship models, remain unknown. We benchmarked 14 LLM configurations, seven Chinese-developed ("domestic") and seven international models, on 1,000 whole-body 18F-FDG PET/CT reports split into an error-injected "junior-docto" arm and a low-residual "finalised" arm (500 each), using a controlled error-injection gold standard. Under each blinded zero-shot prompt, each model flagged six error types and assigned a 1-5 overall score. Two distinct abilities: error-detection macro-F1 (0.356-0.667) and overall-score calibration (ICC[2,1] 0.099-0.627), were weakly and not significantly correlated across models (Spearman {rho} = 0.38, p = 0.18); the dissociation was instead evident in sharp rank reversals, the strongest detector (Claude-Opus-4.8 0.667) calibrating poorly (0.491), while the three best-calibrated models were all domestic (MiMo 0.627, GLM-5 0.612, DeepSeek 0.609). Once the access channel was controlled, domestic and international error detection were statistically indistinguishable ({Delta}macro-F1= -0.011, P = 0.84); domestic models showed consistent but not significant advantages in calibration ({Delta}ICC = +0.142) and Chinese-character-error detection ({Delta}F1 = +0.109), accompanied with large reductions in cost (US$0.09-2.71 vs $0.26-14.5 per 1,000 reports) and on-premise deployability. Re-running two flagships through both agent channels and clean APIs showed that agent channel inflated both detection and calibration (GPT-5.5 {Delta}ICC = +0.098, 95% CI 0.070-0.128), confirming that uncontrolled benchmarks over-credit agent-channel models. Missed-diagnosis detection was the universal weakness (best 0.467) and the one category where the human physicians outperformed every model. Raw detection ability does not guarantee a trustworthy score, and domestic and international models differ by deployment-relevant profile rather than by overall performance rank; both essential distinctions for performing clinical nuclear-medicine QC.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.3%
18.9%
2
Scientific Reports
3612 papers in training set
Top 10%
6.9%
3
PLOS ONE
5266 papers in training set
Top 27%
5.7%
4
The Lancet Digital Health
25 papers in training set
Top 0.1%
5.6%
5
European Journal of Nuclear Medicine and Molecular Imaging
20 papers in training set
Top 0.1%
5.3%
6
Medical Physics
14 papers in training set
Top 0.1%
5.0%
7
GigaScience
212 papers in training set
Top 1%
3.5%
50% of probability mass above
8
Nature Communications
5641 papers in training set
Top 34%
3.3%
9
Journal of the American Medical Informatics Association
71 papers in training set
Top 1%
2.8%
10
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.3%
2.5%
11
Communications Medicine
113 papers in training set
Top 2%
2.2%
12
Imaging Neuroscience
282 papers in training set
Top 2%
2.2%
13
European Radiology
15 papers in training set
Top 0.3%
2.2%
14
Nature Machine Intelligence
70 papers in training set
Top 1%
1.7%
15
Scientific Data
209 papers in training set
Top 2%
1.5%
16
Frontiers in Digital Health
24 papers in training set
Top 0.9%
1.5%
17
PLOS Digital Health
106 papers in training set
Top 3%
1.5%
18
IEEE Access
35 papers in training set
Top 0.8%
1.4%
19
Frontiers in Oncology
103 papers in training set
Top 2%
1.4%
20
iScience
1154 papers in training set
Top 22%
1.4%
21
Communications Biology
993 papers in training set
Top 20%
1.2%
22
PLOS Global Public Health
344 papers in training set
Top 7%
1.2%
23
eBioMedicine
183 papers in training set
Top 4%
1.2%
24
Journal of Cerebral Blood Flow & Metabolism
42 papers in training set
Top 0.5%
1.1%
25
International Journal of Radiation Oncology*Biology*Physics
25 papers in training set
Top 0.4%
1.1%
26
Diagnostics
50 papers in training set
Top 2%
1.1%
27
NeuroImage
903 papers in training set
Top 6%
0.9%
28
PLOS Computational Biology
1863 papers in training set
Top 19%
0.9%
29
Journal of Medical Imaging
11 papers in training set
Top 0.4%
0.9%
30
Science Advances
1243 papers in training set
Top 29%
0.9%