A Neuro-Symbolic Knowledge Graph and Large Language Model Hybrid Architecture for Multi-Modality Mental Health Counseling
Tao, J.; Fenn, N.; Parent, H.; Wu, H.; Arnold, T.; Etue, J.; Chen, E.; Chan, P.
Show abstract
Background: Depression and anxiety are managed largely between clinical visits, yet outpatient care lacks scalable, accountable mechanisms for between-visit support. Large language models converse fluently but fuse clinical reasoning with language generation in one opaque process, so they cannot reliably deliver evidence-based psychotherapy and typically operate outside clinician oversight. Objective: To evaluate C-Mind, a provider-supervised neuro-symbolic system in which a Clinical Knowledge Graph (KG) governs therapeutic decisions for a large language model across eight psychotherapy modalities. Methods: Two simulation regimes addressed eight pre-specified governance questions: a structural validation of KG routing against 117 guideline-anchored vignettes, and a governance battery using progressively disclosing LLM patient agents to evaluate decision traceability, repeatability, provenance auditability, adversarial crisis-detection robustness (277 probes), provider treatment-goal governance, and counselor technique adherence. Crisis detection was additionally validated externally against an independent, clinician-annotated corpus (CRADLE Bench). Results: The KG routed 116/117 vignettes (99.1%) to guideline-appropriate care and detected all 18 high-risk presentations, firing a therapy-suppressing hard halt on 16/18. Adversarial crisis-detection sensitivity was 96.7% and specificity 95.4% (277 probes); on external validation, the system detected 98.5% of 600 dialogues with ongoing suicidal ideation or self-harm at or before the annotator confirming turn. Decisions were 99.1% repeatable, 100% reconstructable per turn, and 100% provenance-auditable across all 354 KG nodes. Provider-set diagnosis, goals, and safety context governed behavior deterministically. Stripped of governance, the same model produced unsolicited clinical monologues on 100% of turns (vs 9% governed) and delivered diagnoses and medication advice the governed system never produced. Conclusions: A neuro-symbolic architecture achieves near-perfect guideline-appropriate routing with a governance profile, traceability, reproducibility, machine-traceable provenance, externally validated crisis detection, and deterministic provider control aligned with requirements for regulated clinical AI.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- New Model, Old Risks? Sociodemographic Bias and Adversarial Hallucinations Vulnerability in GPT-5 92%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 92%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
Similar papers in this journal
Similar papers in this journal
- AI chatbots not yet ready for clinical use 91%
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 90%
- Precision Digital Intervention for Depression Based on Social Rhythm Principles Adds Significantly to Outpatient Treatment 90%
Similar papers in this journal
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 91%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 90%
- Clinical trial emulation can identify new opportunities to enhance the regulation of drug safety in pregnancy 88%
Similar papers in this journal
- Applications of Large Language Models in Psychiatry: A Systematic Review 89%
- Leveraging Reddit data for Context-enhanced Synthetic Health Data Generation to Identify Low Self Esteem 89%
- Interactive Psychometrics for Autism with the Human Dynamic Clamp: Interpersonal Synchrony from Sensory-motor to Socio-cognitive Domains 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.