The Cognitive Safety Net: Comparing Human and AI Diagnostic Reasoning during Complex Clinical Situations.
Audion, A.; Henkeme, M.; Balanca, B.; Lilot, M.; Rimmele, T.; Abaakil, I.; Cejka, J.-C.
Show abstract
BackgroundDiagnostic error in high-stakes clinical environments remains a significant cause of preventable harm. While a new generation of customisable digital cognitive aids (cDCAs) has shown a capacity to improve performance, achieve robust competence, and double learning retention, the potential for artificial intelligence (AI) to augment the foundational, anticipatory reasoning that precedes action is not well understood. This study aims to compare the diagnostic reasoning strategies of experienced anaesthesiology residents with those of a large language model (LLM) during a simulated, complex and realistic anaesthesiology scenario. MethodsWe conducted a comparative analysis within a high-fidelity simulation randomised controlled trial (Anticipamax, NCT06487208). Thirty-four experienced anaesthesiology residents and a conversational LLM (ChatGPT-4) managed a perioperative shock of deliberately multifactorial aetiology. Diagnostic lotteries--sets of hypotheses with assigned plausibility scores--were collected before and after the simulation. We implemented a novel analytical framework based on the social choice Condorcet method, to rank not only individual hypotheses but also to compare the complete diagnostic strategies as the case evolved. ResultsThe AI and residents demonstrated distinct reasoning profiles. Initially, the AI produced an exhaustive, non-hierarchical analysis, correctly identifying septic shock among its top, similarly-scored hypotheses. Residents, in contrast, employed a pragmatic, focused strategy, prioritising immediate surgical risks and unanimously identifying an experience-based risk (gas embolism) that the AI systematically overlooked, and consistently reserved a portion of their reasoning for uncertainty, termed Place for Doubt. After the clinical evolution, both converged on septic shock. A complex scrutiny analysis of the overall strategies revealed that the residents focused and adaptive reasoning was consistently ranked as strategically superior to the AIs exhaustive but diluted approach. ConclusionsAI demonstrates a powerful capacity for broad diagnostic anticipation, acting as a potential safeguard against premature diagnostic closure. Experienced residents exhibit a strategically superior reasoning process in its focus and adaptation. Our findings support a powerful synergy where the AI serves as a Cognitive Safety Net to augment, not replace, the contextualised judgment of the human practitioner. Research in ContextO_ST_ABSWhat is already known on this topicC_ST_ABSO_LIHuman error in healthcare is a global prominent cause of death. C_LIO_LI Traditional cognitive support tools (e.g., paper checklists) have been shown to improve technical skills during medical crises, but their impact on non-technical skills is limited and their clinical adoption remains low. C_LIO_LIA new generation of customisable digital cognitive aids (cDCAs) can significantly improve both technical and non-technical performance, fostering better team management and crisis resolution. C_LIO_LIInformation on how clinicians deliver the best anticipatory clinical reasoning is scarce. C_LIO_LIRecent work comparing machine-learning models to clinicians in trauma triage found comparable accuracy but only moderate agreement, suggesting a collaborative paradigm and motivating deeper analyses of the reasoning process itself. C_LIO_LIHowever, a critical gap remains in understanding the underlying nature of the diagnostic reasoning strategies that lead to these outcomes. The how of human and AI reasoning, especially in dynamic, anticipatory clinical tasks, is not well understood. C_LI What this study addsO_LIThis is the first study to directly compare in action the diagnostic reasoning strategies of clinicians and a large language model (AI). C_LIO_LIIt introduces a novel analytical framework based on the Condorcet social choice method to move beyond simple performance scores and rigorously model and rank the overall quality of diagnostic strategies in a simulated daily complex situation. C_LIO_LIThe findings support a model of human-AI complementarity, where the AI excels at broad, exhaustive analysis, while clinicians demonstrate a superior, focused, and adaptive strategic reasoning, suggesting the humans role as a meta-cognitive supervisor of AI-driven exhaustive but diluted insights. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 93%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 91%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 91%
Similar papers in this journal
- Communicating personalised risks from COVID-19: guidelines from an empirical study 90%
- Discrepancy review: A feasibility study of a novel peer review intervention to reduce undisclosed discrepancies between registrations and publications 88%
- Getting stuck in a rut as an emergent feature of a dynamic decision-making system 88%
Similar papers in this journal
- Using explainable machine learning to identify patients at risk of reattendance at discharge from emergency departments 94%
- Machine learning approach to dynamic risk modeling of mortality in COVID-19: a UK Biobank study 92%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 92%
Similar papers in this journal
- Modelling Palliative and End of Life resource requirements during COVID-19: implications for quality care 91%
- How and why do Quality Circles work for General Practitioners - a realist approach 91%
- Development and validation of automated computer aided-risk score for predicting the risk of in-hospital mortality using first electronically recorded blood test results and vital signs for COVID-19 hospital admissions: a retrospective development and validation study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.