Context-Dependent FHIR Serialisation Strategies for Clinical LLM Deployment: A Multi-Layer Benchmark on UK Core Data
Chong, J.
Show abstract
The choice of FHIR-to-text serialisation format significantly impacts clinical LLM quality (Kruskal-Wallis H=163.86, p<10^-33, delta=0.24 on a 5-point scale), yet remains unstudied as a clinical deployment variable. We present FHIRBench-UK, evaluating five large language models across six serialisation formats and three clinical tasks on 100 UK Core FHIR patient bundles (18,000 scored prompts across clean and perturbed cohorts). Our findings converge with independent work on open-weight models (Pator, 2026). The optimal format is context-dependent: raw_json dominates for clinical QA, hybrid_adaptive for clinical reasoning, and structured_markdown for summarisation. In 58% of model-task-complexity scenarios, raw_json is suboptimal. Model capability moderates format sensitivity: Claude Sonnet 4.5 shows 0.10-point sensitivity versus Llama 3.3's 0.39, making adaptive serialisation most valuable for budget-constrained deployments using mid-tier models. All findings replicate under clinically realistic data perturbation. The study additionally confirms a complete ranking inversion between token-level F1 and clinical quality (rho=-0.90), replicating US findings across UK Core profiles. We recommend task-aware serialisation routing as a zero-cost quality intervention for NHS FHIR-based LLM deployments.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 92%
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 91%
- AI chatbots not yet ready for clinical use 90%
Similar papers in this journal
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 94%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 93%
- Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review 92%
Similar papers in this journal
- Toward Trustworthy Chatbots: A Protocol for Red Teaming for Health Related Conversations 93%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 92%
- Large Language Models for Zero-Shot Procedure Extraction in Orthopedic Surgery: A Comparative Evaluation 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.