Back

Semantic Encoding in Medical LLMs for Vocabulary Standardisation

Mainwood, S.; Bhandari, A.; Tyagi, S.

2025-06-17 health informatics
10.1101/2025.06.16.25329716 medRxiv
Show abstract

High-quality, standardised medical data availability remains a bot-tleneck for digital health and AI model development. A major hurdle is translating noisy free text into controlled clinical vocabularies, aiming for harmonisation and interoperability, especially when source datasets are inconsistent or incomplete. We bench-mark domain-specific encoder models against general LLMs for semantic-embedding retrieval using minimal vocabulary building blocks and test several prompt techniques. We also try prompt augmentation with LLM-generated differential definitions. We tested these prompts on open-source Llama and medically fine-tuned Llama models to steer their alignment toward accurate concept assignment across multiple prompt formats. Domain-tuned models consistently outperform general models of the same size in retrieval and generative tasks. However, performance is sensitive to prompt design and model size, and the benefits of adding LLM-generated context are inconsistent. While newer, larger foundation models are closing the gap, todays lightweight open-source generative LLMs lack the stability and embedded clinical knowledge needed for deterministic vocabulary standardisation.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.