Entity-centric evaluation of large language model responses for medical question-answering tasks
Liu, Y.; Kolachalama, V. B.
Show abstract
ObjectiveDevelop a metric for evaluating the clinical alignment and informativeness of large language model (LLM)-generated responses in medical question-answering (QA) tasks. Materials and methodsWe propose EntQA, an entity-centric metric that extracts biomedical entities from patient backgrounds, diagnostic questions and LLM responses using a biomedical named entity recognition model, followed by de-duplication and semantic/lexical matching with thresholds. We computed recall-style coverage scores to quantify entity retention and detect omissions without external resources. We evaluated EntQA on five benchmarks using seven Qwen 2.5 Instruct models (0.5B-72B parameters), comparing it to baselines via Spearman/Kendall correlations with model accuracy at group level, point-biserial correlations at case level, and Spearman correlations with model scaling. ResultsEntQA demonstrated consistent positive alignments with accuracy (group-level Spearman up to 0.9286; case-level point-biserial up to 0.0926) and model scaling (Spearman up to 0.252), outperforming baselines which often showed negative or inconsistent correlations (e.g., BERTScore Spearman -0.9286 with accuracy). ConclusionEntQA offers a scalable, interpretable evaluation for LLM medical QA, outperforming traditional metrics in capturing clinical fidelity and supporting trustworthy healthcare AI through applications in fact-checking and model refinement.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
Similar papers in this journal
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 95%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 95%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 94%
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 94%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 91%
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 96%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 94%
- Automating Evaluation of LLM-generated Responses to Patient Questions about Rare Diseases 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.