Multi-model LLM assessment of Quality Control Circlemethodological quality: a designed-anchor reliabilitystudy
LIn, H.; Lyu, J.
Show abstract
BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Incorporating Preprints in Systematic Reviews: A Preliminary Study of a Novel Method for Rapid Evidence Synthesis 93%
- Beyond Metrics to Methods: A Scoping Review of Large Language Models for Detection of Social Drivers of Health in Clinical Notes 92%
- Sociodemographic Bias in Large Language Model Clinical Trial Screening 91%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 93%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 92%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 92%
Similar papers in this journal
- Fast and frugal decision tree for the rapid critical appraisal of systematic reviews 93%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 93%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 93%
Similar papers in this journal
- Agreeability testing of AMSTAR-PF, a tool for quality appraisal of systematic reviews of prognostic factor studies 93%
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 93%
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.