Back

The Cognitive Safety Net: Comparing Human and AI Diagnostic Reasoning during Complex Clinical Situations.

Audion, A.; Henkeme, M.; Balanca, B.; Lilot, M.; Rimmele, T.; Abaakil, I.; Cejka, J.-C.

2025-10-07 health systems and quality improvement
10.1101/2025.10.06.25335641 medRxiv
Show abstract

BackgroundDiagnostic error in high-stakes clinical environments remains a significant cause of preventable harm. While a new generation of customisable digital cognitive aids (cDCAs) has shown a capacity to improve performance, achieve robust competence, and double learning retention, the potential for artificial intelligence (AI) to augment the foundational, anticipatory reasoning that precedes action is not well understood. This study aims to compare the diagnostic reasoning strategies of experienced anaesthesiology residents with those of a large language model (LLM) during a simulated, complex and realistic anaesthesiology scenario. MethodsWe conducted a comparative analysis within a high-fidelity simulation randomised controlled trial (Anticipamax, NCT06487208). Thirty-four experienced anaesthesiology residents and a conversational LLM (ChatGPT-4) managed a perioperative shock of deliberately multifactorial aetiology. Diagnostic lotteries--sets of hypotheses with assigned plausibility scores--were collected before and after the simulation. We implemented a novel analytical framework based on the social choice Condorcet method, to rank not only individual hypotheses but also to compare the complete diagnostic strategies as the case evolved. ResultsThe AI and residents demonstrated distinct reasoning profiles. Initially, the AI produced an exhaustive, non-hierarchical analysis, correctly identifying septic shock among its top, similarly-scored hypotheses. Residents, in contrast, employed a pragmatic, focused strategy, prioritising immediate surgical risks and unanimously identifying an experience-based risk (gas embolism) that the AI systematically overlooked, and consistently reserved a portion of their reasoning for uncertainty, termed Place for Doubt. After the clinical evolution, both converged on septic shock. A complex scrutiny analysis of the overall strategies revealed that the residents focused and adaptive reasoning was consistently ranked as strategically superior to the AIs exhaustive but diluted approach. ConclusionsAI demonstrates a powerful capacity for broad diagnostic anticipation, acting as a potential safeguard against premature diagnostic closure. Experienced residents exhibit a strategically superior reasoning process in its focus and adaptation. Our findings support a powerful synergy where the AI serves as a Cognitive Safety Net to augment, not replace, the contextualised judgment of the human practitioner. Research in ContextO_ST_ABSWhat is already known on this topicC_ST_ABSO_LIHuman error in healthcare is a global prominent cause of death. C_LIO_LI Traditional cognitive support tools (e.g., paper checklists) have been shown to improve technical skills during medical crises, but their impact on non-technical skills is limited and their clinical adoption remains low. C_LIO_LIA new generation of customisable digital cognitive aids (cDCAs) can significantly improve both technical and non-technical performance, fostering better team management and crisis resolution. C_LIO_LIInformation on how clinicians deliver the best anticipatory clinical reasoning is scarce. C_LIO_LIRecent work comparing machine-learning models to clinicians in trauma triage found comparable accuracy but only moderate agreement, suggesting a collaborative paradigm and motivating deeper analyses of the reasoning process itself. C_LIO_LIHowever, a critical gap remains in understanding the underlying nature of the diagnostic reasoning strategies that lead to these outcomes. The how of human and AI reasoning, especially in dynamic, anticipatory clinical tasks, is not well understood. C_LI What this study addsO_LIThis is the first study to directly compare in action the diagnostic reasoning strategies of clinicians and a large language model (AI). C_LIO_LIIt introduces a novel analytical framework based on the Condorcet social choice method to move beyond simple performance scores and rigorously model and rank the overall quality of diagnostic strategies in a simulated daily complex situation. C_LIO_LIThe findings support a model of human-AI complementarity, where the AI excels at broad, exhaustive analysis, while clinicians demonstrate a superior, focused, and adaptive strategic reasoning, suggesting the humans role as a meta-cognitive supervisor of AI-driven exhaustive but diluted insights. C_LI

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.