Theory of Mind Imitation by LLMs for Physician-Like Human Evaluation
Awasthi, R.; Mishra, S.; Raghu, C.; Moises, A.; Atreja, A.; Mahapatra, D.; Singh, N.; Khanna, A. K.; Cywinski, J. B.; Maheshwari, K.; Papay, F. A.; Mathur, P.
Show abstract
Aligning the Theory of Mind (ToM) capabilities of Large Language Models (LLMs) with human cognitive processes enables them to imitate physician behavior. This study evaluates LLMs abilities such as Belief and Knowledge, Reasoning and Problem-Solving, Communication and Language Skills, Emotional and Social Intelligence, Self-Awareness, and Metacognition in performing human-like evaluations of Foundation Models. We used a dataset composed of clinical questions, reference answers, and LLM-generated responses based on guidelines for the prevention of heart disease. Comparing GPT-4 to human experts across ToM abilities, we found the highest Emotional and Social Intelligence agreement using the Brennan-Prediger coefficient. This study contributes to a deeper understanding of LLMs cognitive capabilities and highlights their potential role in augmenting or complementing human clinical assessments.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 92%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
- Simulated Misuse of Large Language Models and Clinical Credit Systems 92%
Similar papers in this journal
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 92%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 92%
Similar papers in this journal
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 93%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 92%
- User Experience Evaluation of Cogscreen for Screening Mild Cognitive Impairment: Formative and Summative Evaluation 92%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 91%
- AI-Generated Clinical Summaries: Errors and Susceptibility to Speech and Speaker Variability 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.