CardiacGPT: A Real-Time AI Assistant for Intraoperative Guidance and Postoperative Decision Support in Cardiac Surgery
Ramezani, A.; Moradmand, A.; Nezami, F. R.
Show abstract
BackgroundCardiac surgery is one of the most complex and high-stakes areas of medicine, where intraoperative decisions must be made within seconds and incomplete information can compromise outcomes. Traditional risk scores and rule-based decision support tools provide limited real-time guidance and rarely integrate the unstructured data streams available during surgery. Recent advances in large language models (LLMs) such as OpenAIs GPT-5 and Anthropics Claude 3.5 family have demonstrated state-of-the-art reasoning, summarization, and clinical dialogue capabilities. However, their safety and trustworthiness in surgical settings remain untested. ObjectiveTo evaluate the feasibility and clinician trustworthiness of CardiacGPT, a real-time AI assistant that leverages the newest-generation LLMs for intraoperative guidance and postoperative decision support. MethodsWe retrospectively analyzed 500 de-identified cardiac surgery cases from Brigham and Womens Hospital, including CABG, valve, and combined procedures. Structured EHR variables, intraoperative monitoring, and operative notes were formatted into standardized prompts and processed through four cutting-edge models: OpenAI GPT-5, Anthropic Claude 3.5 Opus, Claude 3.5 Sonnet, and Claude 3.5 Haiku. Outputs were presented via a blinded Bidding App to attending cardiac surgeons and ICU clinicians, who scored trust and clinical relevance on a 5-point Likert scale. The primary outcome was the proportion of high-trust ratings (score > 4); secondary outcomes included mean trust scores, variance, and inter-rater reliability. ResultsAcross 2,000 evaluations, GPT-5 and Claude 3.5 Opus achieved the highest mean trust scores (4.83 and 4.79, respectively), each exceeding 98% high-trust ratings. Claude 3.5 Sonnet performed moderately (mean 3.9, 74% high-trust), while Claude 3.5 Haiku produced less context-specific recommendations (mean 3.6, 66% high-trust). Inter-rater reliability was excellent, with ICC(2,1) = 0.91 (95% CI 0.88-0.94), confirming strong agreement among reviewers. Qualitative analysis showed that GPT-5 and Claude 3.5 Opus generated actionable and context-aware outputs, whereas smaller models often produced generic or incomplete guidance. ConclusionsCardiacGPT, powered by the newest LLMs (GPT-5 and Claude 3.5 series), demonstrated feasibility and exceptionally high clinician trust across 500 real-world surgical cases. This is the first blinded, multi-model evaluation of next-generation LLMs for cardiac surgery. While outcome-based prospective trials are still required, these results establish CardiacGPT as a promising real-time co-pilot for cardiac surgeons, with the potential to reduce cognitive load, standardize intraoperative communication, and improve postoperative planning.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 93%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
Similar papers in this journal
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 93%
- Evaluating algorithmic fairness in the presence of clinical guidelines: the case of atherosclerotic cardiovascular disease risk estimation 90%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 90%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 95%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 93%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 92%
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 93%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 93%
Similar papers in this journal
- Decision support system to evaluate VENTilation in the Acute Respiratory Distress Syndrome (DeVENT study) – Trial Protocol 91%
- Electronic early notification of sepsis in hospitalized ward patients: a study protocol for a stepped-wedge cluster randomized controlled trial 90%
- Supporting Reanalysis and Reuse of Clinical Trial Data: A Case Study 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.