Semantic Encoding in Medical LLMs for Vocabulary Standardisation
Mainwood, S.; Bhandari, A.; Tyagi, S.
Show abstract
High-quality, standardised medical data availability remains a bot-tleneck for digital health and AI model development. A major hurdle is translating noisy free text into controlled clinical vocabularies, aiming for harmonisation and interoperability, especially when source datasets are inconsistent or incomplete. We bench-mark domain-specific encoder models against general LLMs for semantic-embedding retrieval using minimal vocabulary building blocks and test several prompt techniques. We also try prompt augmentation with LLM-generated differential definitions. We tested these prompts on open-source Llama and medically fine-tuned Llama models to steer their alignment toward accurate concept assignment across multiple prompt formats. Domain-tuned models consistently outperform general models of the same size in retrieval and generative tasks. However, performance is sensitive to prompt design and model size, and the benefits of adding LLM-generated context are inconsistent. While newer, larger foundation models are closing the gap, todays lightweight open-source generative LLMs lack the stability and embedded clinical knowledge needed for deterministic vocabulary standardisation.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 96%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 95%
- Towards semantic interoperability: finding and repairing hidden contradictions in biomedical ontologies 94%
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 97%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 94%
Similar papers in this journal
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 94%
- Comparing neural language models for medical concept representation and patient trajectory prediction 93%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.