Back

Large Language Model Performance in UK Advice & Guidance: A Pilot Study in Neurology

Healy, J.; Marvasti, A.; Wallace, D.; Baheerathan, A.; Ghosh, A.; Kossoff, J.; Thio, S.; Balaratnam, M.; Haider, S.; Ellershaw, S.; Dobson, R.

2026-05-18 neurology
10.64898/2026.05.13.26353081 medRxiv
Show abstract

Background: Large language models (LLMs) demonstrate strong performance in controlled medical environments such as multiple choice exams, but their utility in real-world clinical workflows remains unproven. The NHS Advice & Guidance (A&G) service, where Primary Care clinicians can submit text-based queries to specialists, provides an environment for evaluating the clinical performance of LLMs as a specialist. Methods: We compared responses from MedGemma 4B-IT, an open-weight model deployed locally on hospital infrastructure, against specialist neurologist responses across 50 adult neurology A&G cases from University College London Hospital. Two neurologists and two GPs rated 80 blinded and 20 unblinded responses for outcome, safety, efficacy, and feasibility using standardised criteria; outcome was a binary correct/incorrect, while other domains were scored 1-5. Inter-rater reliability was assessed using intraclass correlation coefficients. Results: Although there were no statistically significant differences between blinded specialist neurologists and LLM responses across any domain (outcome: 84% vs 82%, p=0.67; safety: 3.98 vs 4.02, p=0.85; efficacy: 4.06 vs 3.98, p=0.61; feasibility: 4.39 vs 4.20, p=0.45), 10% of LLM responses received concerning scores ([≤]2 average score) compared to 0% of human responses, indicating potentially clinically important tail risk. Furthermore, unblinded results showed a preference for human responses, with human ratings being preferred across all domains. Only 51% of binary outcomes had unanimous agreement and inter-rater agreement was moderate across other domains (ICC 0.50-0.52). Conclusions: In this pilot study, aggregate scores between blinded human and LLM responses were similar, and no statistically significant differences were detected in this exploratory sample. However, aggregate metrics masked clinically important edge-case failures in LLM responses. Pronounced inter-rater variability and the potential impact of LLM/human syntax on blinded rater judgements highlight the challenges in establishing robust evaluation frameworks for clinical LLM deployment

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
BMJ Health & Care Informatics
15 papers in training set
Top 0.1%
9.0%
2
npj Digital Medicine
118 papers in training set
Top 0.7%
7.9%
3
PLOS ONE
5266 papers in training set
Top 21%
7.9%
4
PLOS Digital Health
106 papers in training set
Top 0.8%
6.8%
5
Behavior Research Methods
30 papers in training set
Top 0.1%
5.5%
6
Frontiers in Neurology
102 papers in training set
Top 0.9%
4.1%
7
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.6%
3.3%
8
Frontiers in Digital Health
24 papers in training set
Top 0.4%
3.2%
9
The Lancet Digital Health
25 papers in training set
Top 0.2%
2.7%
50% of probability mass above
10
Journal of the American Medical Informatics Association
71 papers in training set
Top 1%
2.5%
11
Pharmacoepidemiology and Drug Safety
18 papers in training set
Top 0.2%
2.5%
12
Journal of NeuroEngineering and Rehabilitation
36 papers in training set
Top 0.4%
2.1%
13
BMC Medicine
176 papers in training set
Top 2%
2.1%
14
GigaScience
212 papers in training set
Top 2%
1.9%
15
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.9%
16
BMJ Open
601 papers in training set
Top 9%
1.7%
17
Communications Medicine
113 papers in training set
Top 2%
1.7%
18
Scientific Reports
3612 papers in training set
Top 60%
1.4%
19
Orphanet Journal of Rare Diseases
21 papers in training set
Top 0.4%
1.3%
20
BMC Neurology
14 papers in training set
Top 0.4%
1.3%
21
Journal of Neurology, Neurosurgery & Psychiatry
30 papers in training set
Top 0.5%
1.3%
22
BJGP Open
13 papers in training set
Top 0.4%
1.1%
23
Emergency Medicine Journal
21 papers in training set
Top 0.3%
1.1%
24
Epilepsia
56 papers in training set
Top 0.5%
1.1%
25
iScience
1154 papers in training set
Top 29%
1.0%
26
Brain Communications
166 papers in training set
Top 3%
1.0%
27
JMIR Medical Informatics
18 papers in training set
Top 0.8%
1.0%
28
Artificial Intelligence in Medicine
17 papers in training set
Top 0.7%
0.9%
29
JMIR Formative Research
33 papers in training set
Top 1%
0.9%
30
Annals of Clinical and Translational Neurology
34 papers in training set
Top 0.9%
0.9%