Do Large Language Models Have a Personality? A Psychometric Evaluation with Implications for Clinical Medicine and Mental Health AI
Heston, T. F.; Gillette, J.
Show abstract
IntroductionLarge language models (LLMs) are increasingly used in clinical medicine to provide emotional support, deliver cognitive-behavioral therapy, and assist in triage and diagnosis. However, as LLMs are integrated into mental health applications, assessing their inherent personality traits and evaluating their divergence from expected neutrality is essential. This study characterizes the personality profiles exhibited by LLMs using two validated frameworks: the Open Extended Jungian Type Scales (OEJTS) and the Big Five Personality Test. MethodsFour leading LLMs publicly available in 2024 [ChatGPT-3.5 (OpenAI), Gemini Advanced (Google), Claude 3 Opus (Anthropic), and Grok-Regular Mode (X)] were evaluated across both psychometric instruments. A one-way multivariate analysis of variance (MANOVA) was performed to assess inter-model differences in personality profiles. ResultsMANOVA demonstrated statistically significant differences across models in typological and dimensional personality traits (Wilks Lamda = 0.115, p < 0.001). OEJTS results showed ChatGPT-3.5 most often classified as ENTJ and Claude 3 Opus consistently as INTJ, while Gemini Advanced and Grok-Regular leaned toward INFJ. On the Big Five Personality Test, Gemini scored markedly lower on agreeableness and conscientiousness, while Claude scored highest on conscientiousness and emotional stability. Grok-Regular exhibited high openness but more variability in stability. Effect sizes ranged from moderate to large across traits. ConclusionDistinct personality profiles are consistently expressed across different LLMs, even in unprompted conditions. Given the increasing integration of LLMs into clinical workflows, these findings underscore the need for formal personality evaluation and oversight involving mental health professionals before deployment.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Can Psychedelic Use Benefit Meditation Practice? Examining Individual, Psychedelic, and Meditation-Related Factors 92%
- The Relationship between Latent Inhibition, Divergent Thinking, and Eyewitness Memory: A Study on Attention to Irrelevant Stimuli 92%
- Cross-Cultural Validation of Acculturation Measures: Expanding the East Asian Acculturation Framework for Global Applicability 92%
Similar papers in this journal
- Exploring patterns of ongoing thought under naturalistic and task-based conditions 92%
- Can Hypnotic Susceptibility be Explained by Bifactor Models? Structural Equation Modeling of the Harvard Group Scale of Hypnotic Susceptibility - Form A 92%
- Sense of Agency for Mental Actions: Insights from a Belief-Based Action-Effect Paradigm 91%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 92%
- The development of a World Health Organization transdiagnostic chatbot intervention for distressed adolescents and young adults 92%
- Remote digital measurement of visual and auditory markers of Major Depressive Disorder severity and treatment response. 88%
Similar papers in this journal
- Social Perception and Interaction Database - a novel tool to study social cognitive processes with point-light displays. 92%
- Co-development of a best practice checklist for mental health data science: A Delphi study 92%
- Passive sensing data predicts stress in university students: A supervised machine learning method for digital phenotyping 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.