Identification and Description of Emotions by Current Large Language Models
Patel, S. C.; Fan, J.
Show abstract
The assertion that artificial intelligence (AI) cannot grasp the complexities of human emotions has been a long-standing debate. However, recent advancements in large language models (LLMs) challenge this notion by demonstrating an increased capacity for understanding and generating human-like text. In this study, we evaluated the empathy levels and the identification and description of emotions by three current language models: Bard, GPT 3.5, and GPT 4. We used the Toronto Alexithymia Scale (TAS-20) and the 60-question Empathy Quotient (EQ-60) questions to prompt these models and score the responses. The models performance was contrasted with human benchmarks of neurotypical controls and clinical populations. We found that the less sophisticated models (Bard and GPT 3.5) performed inferiorly on TAS-20, aligning close to alexithymia, a condition with significant difficulties in recognizing, expressing, and describing ones or others experienced emotions. However, GPT 4 achieved performance close to the human level. These results demonstrated that LLMs are comparable in their ability to identify and describe emotions and may be able to surpass humans in their capacity for emotional intelligence. Our novel insights provide alignment research benchmarks and a methodology for aligning AI with human values, leading toward an empathetic AI that mitigates risk.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Investigating the Role of AI Explanations in Lay Individuals’ Comprehension of Radiology Reports: A Metacognition Lense 94%
- The Mastery Rubric for Bioinformatics: supporting design and evaluation of career-spanning education and training 93%
- Proprioceptive accuracy in immersive virtual reality: A developmental perspective 92%
Similar papers in this journal
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 93%
- Can you affect me? The influence of vitality forms on action perception and motor response 92%
- Development and testing of a game-based digital intervention for working memory training in autism spectrum disorder 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.