Towards Richer AI-Assisted Psychotherapy Note-Making and Performance Benchmarking
Adhikary, P. K.; Singh, S.; Singh, S.; Sharma, P.; Soni, P.; Choudhary, R.; Saxena, C.; Chauhan, P.; Gupta, S. K.; Deb, K. S.; Singh, S. M.; Chakraborty, T.
Show abstract
Psychotherapy note-making is crucial for effective patient care. However, traditional formats such as SOAP (Subjective, Objective, Assessment, and Plan) and BIRP (Behavior, Intervention, Response, and Plan) often fail to capture the nuanced complexities of therapeutic sessions, as they primarily focus on surface-level details and lack a comprehensive understanding of the patients history, mental status, and therapeutic process. While recent advances in Artificial Intelligence (AI) and Large Language Models (LLMs) show promise in clinical documentation, their application in psychotherapy note summarisation remains unexplored. We present iCARE (identifiers, Chief Concerns and Clinical History, Assessment and Analysis, Risk and Crisis, Engagement and Next Steps), a comprehensive framework for AI-assisted psychotherapy documentation that addresses these limitations. iCARE comprises of 17 clinically relevant aspects, developed collaboratively with mental health professionals, and aligned with established guidelines. We further introduce PATH (Psychotherapy Aspects and Treatment History summary), a novel dataset of annotated therapy sessions. Through extensive benchmarking with 11 LLMs, including both open and closed-source models, we evaluate their performance across different note-taking aspects using automatic and human evaluation metrics. Our results show that closed-source models like Gemini Pro and GPT4o-mini excel in various aspects, with Gemini Pro achieving superior human evaluation scores. Notably, all models struggle with temporal reasoning and complex therapeutic interpretations. The findings suggest that current LLMs can assist in basic documentation but require improvements in handling longitudinal therapeutic relationships and aspects that require deeper clinical understanding and interpretative reasoning. This work advances mental health care documentation while emphasising the need for continued clinical expertise in psychotherapy note summarisation.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Applications of Large Language Models in Psychiatry: A Systematic Review 92%
- Computational Psychiatry Research Map (CPSYMAP): a New Database for Visualizing Research Papers 91%
- Development of Goal Management Training + (GMT + ) for Methamphetamine Use Disorder Through Collaborative Design: A Process Description 91%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 92%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 95%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 91%
- The development of a World Health Organization transdiagnostic chatbot intervention for distressed adolescents and young adults 91%
Similar papers in this journal
- Psychotherapies and Psychological Support for Individuals Facing Psychological Distress during the COVID-19 Pandemic: A Scoping Review 92%
- The Benefits and Harms of Open Notes in Mental Health: A Delphi Survey of International Experts 91%
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.