Back

On the robustness of medical term representations in locally deployable language models

Auger, S. D.; Graham, N. S. N.; Scott, G.

2026-02-26 health informatics
10.64898/2026.02.24.26346972 medRxiv
Show abstract

Structured AbstractO_ST_ABSBackgroundC_ST_ABSHosting large language models (LLMs) on-premises can secure patient data but requires compact architectures to function on standard hardware. The impact of such constraints on the robustness of their representations for medical terminology is important for clinical AI safety but poorly understood. The statistical nature of LLM training inherently limits the representation of terms with low societal prominence or lexical frequency, and high ambiguity. MethodsWe assessed 15 open-weights LLMs (4B-120B) for their representational robustness of 250 neurological terms. Neurology was chosen for its strict hierarchical and anatomical terminology. A terms representation was deemed robust only if the model correctly navigated four tests, verifying valid links against distractors and reverse associations. We examined associations between representational robustness and model size, medical fine-tuning, and five terminological subdomains (localisation, clinical features, investigations, diagnoses, and treatments). We assessed term difficulty using the semantic complexity index (SCI), a novel composite integrating societal prominence, lexical frequency, and ambiguity. ResultsRepresentational robustness followed a log-linear scaling law relative to model size (r=0.736, p=0.002). Medical fine-tuning yielded no benefit for 4B models, but significantly improved larger 27B model performance, with rate of robust representations rising from 38.2% to 62.6% (p<0.0001). While most local-LLMs performance degraded sharply with increasing SCI values, GPT-OSS 20B and 120B maintained complexity invariance (with <20% decline from lowest to highest complexity terms). Notably, the general-purpose 20B GPT-OSS model outperformed larger and medically fine-tuned counterparts. Robustness varied by subdomain (F=4.69, p=0.003), with diagnoses (73.8%) scoring significantly higher than localisation (47.9%, p=0.004) and clinical features (52.1%, p=0.02). ConclusionsWhile representational robustness broadly follows model size scaling laws, neither model size nor fine-tuning guarantees clinical reliability. Since performance fluctuates with terminological complexity and subdomain, safe deployment requires validating representational robustness for specific use cases rather than assuming larger models handle medical language safely. 1-2 Sentence DescriptionThis study shows that model size and medical fine-tuning are not reliable indicators of clinical robustness across 15 locally-deployable LLMs. Because performance varies significantly by terminological complexity and subdomain, safe application requires validation methods that account for these factors.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.