Back

Medical concept understanding in large language models is fragmented

Deng, L.; Chen, L.; Liu, M.

2026-03-05 health informatics
10.64898/2026.03.03.26347552 medRxiv
Show abstract

Large language models (LLMs) perform strongly across a wide range of medical applications, yet it remains unclear whether such success reflects genuine understanding of medical concepts. We present an ontology-grounded, concept-centered evaluation of medical concept understanding in LLMs. Using 6,252 phenotype concepts from Human Phenotype Ontology, we decompose concept understanding into three core dimensions--concept identity, concept hierarchy, and concept meaning--and design corresponding benchmarks for each dimension. Across a representative set of contemporary LLMs, best-performing models achieve high accuracy on concept identity (90.6%) and hierarchy (83.8%), but lower performance on concept meaning (72.6%). Concept-level analysis reveals substantial fragmentation in LLM understanding: only 57.7% of concepts are consistently understood across all three dimensions, while 41.3% show partial understanding and 1.1% are not captured in any dimension. These results demonstrate that strong application-level performance of LLMs can mask fundamental gaps in concept-level understanding, highlighting the necessity for ontology-grounded evaluation in medical AI.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.