Back

Theory of Mind Imitation by LLMs for Physician-Like Human Evaluation

Awasthi, R.; Mishra, S.; Raghu, C.; Moises, A.; Atreja, A.; Mahapatra, D.; Singh, N.; Khanna, A. K.; Cywinski, J. B.; Maheshwari, K.; Papay, F. A.; Mathur, P.

2025-03-05 health informatics
10.1101/2025.03.01.25323142 medRxiv
Show abstract

Aligning the Theory of Mind (ToM) capabilities of Large Language Models (LLMs) with human cognitive processes enables them to imitate physician behavior. This study evaluates LLMs abilities such as Belief and Knowledge, Reasoning and Problem-Solving, Communication and Language Skills, Emotional and Social Intelligence, Self-Awareness, and Metacognition in performing human-like evaluations of Foundation Models. We used a dataset composed of clinical questions, reference answers, and LLM-generated responses based on guidelines for the prevention of heart disease. Comparing GPT-4 to human experts across ToM abilities, we found the highest Emotional and Social Intelligence agreement using the Brennan-Prediger coefficient. This study contributes to a deeper understanding of LLMs cognitive capabilities and highlights their potential role in augmenting or complementing human clinical assessments.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.