Back

Combined values alignment and epistemic verification prevent delusional reinforcement in conversational AI agents

Carrano, A.; Patel, M. S.; Hartono, S.; Ekker, S. C.

2026-06-02 health informatics
10.64898/2026.05.29.26354389 medRxiv
Show abstract

Conversational AI is being deployed into medical decision support, mental-health triage, and social companionship, where reinforcement of a user's false or delusional belief can cause direct harm. Most deployed safety techniques are evaluated for factual accuracy in isolation; the question of whether they protect against belief-level harm, and whether layered architectures behave additively or synergistically, has not been answered empirically. We compared four configurations of the same underlying model: a bare language model (condition A); an explicit values constraint we call the First Law architecture (condition B); a real-time epistemic verification layer called Aletheia (condition C); and the complete architecture combining all components together (condition D). Across 156 scored responses spanning 39 probe items in four belief-harm domains, condition A only passed 3 of 36 main-battery probes (8.3%; 95% CI 1.8 to 22.5%) under triple-blind human consensus rating demonstrating the core limitations of unmodified LLM deployments. In contrast, the three safety architectures (B-D) passed at least 97% of items (Fisher's exact, P < 0.001 versus A). On a synergy battery designed to test items at the intersection of value- and epistemic-domain failures (16 scored items, AI-rated), only the complete architecture passed every item; single-layer conditions failed on 7 of 16 items (43.8%) where neither values constraint nor verification was individually sufficient. Linear mixed-effects modelling of three-turn emotional escalation gave a slope of -1.00 points per turn for the values-only condition (t = -6.20) and -0.75 points per turn for the verification-only condition (t = -4.65); the complete architecture was flat at {beta} = 0.00. We describe a mechanistic failure of single-layer verification we call bot-validates-kernel-endorses-inference, in which accurate confirmation of a true factual element embedded in a delusional claim transfers epistemic authority to the surrounding false inference. Values alignment and factual verification address different failure modes, and the combined VaaS-Aletheia architecture is what produces stable protection across emotional escalation in conversational settings. The complete architecture evaluated here represents evidence-based specification for safer deployment of AI in high-stakes advisory contexts and serves as a benchmark against which future safety architectures can be compared.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Nature Medicine
125 papers in training set
Top 0.1%
26.6%
2
npj Digital Medicine
118 papers in training set
Top 0.6%
9.7%
3
PLOS ONE
5266 papers in training set
Top 22%
7.9%
4
Philosophical Transactions of the Royal Society B
51 papers in training set
Top 0.1%
6.7%
50% of probability mass above
5
Scientific Reports
3612 papers in training set
Top 11%
6.7%
6
Nature Communications
5641 papers in training set
Top 31%
4.3%
7
iScience
1154 papers in training set
Top 5%
3.5%
8
Patterns
78 papers in training set
Top 0.8%
2.4%
9
Nature Human Behaviour
95 papers in training set
Top 0.9%
2.1%
10
eLife
5828 papers in training set
Top 60%
1.1%
11
Royal Society Open Science
214 papers in training set
Top 5%
1.1%
12
BMC Medicine
176 papers in training set
Top 4%
1.0%
13
Med
39 papers in training set
Top 0.5%
1.0%
14
Nature
645 papers in training set
Top 9%
1.0%
15
Nature Methods
385 papers in training set
Top 6%
0.8%
16
Science Advances
1243 papers in training set
Top 30%
0.8%
17
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 41%
0.8%
18
Expert Systems with Applications
11 papers in training set
Top 0.4%
0.8%
19
Frontiers in Psychiatry
87 papers in training set
Top 2%
0.8%
20
EMBO Molecular Medicine
95 papers in training set
Top 3%
0.8%
21
Psychological Review
19 papers in training set
Top 0.2%
0.8%
22
Physical Biology
46 papers in training set
Top 0.9%
0.8%
23
PLOS Digital Health
106 papers in training set
Top 4%
0.6%
24
Communications Medicine
113 papers in training set
Top 6%
0.6%