Development and Evaluation of an AI-Assisted, Privacy-Preserving Surgical Risk Calculator
Wolfrath, N.; SenthilKumar, G.; Ramamurthi, A.; Kothari, A. N.
Show abstract
Large language models (LLMs) have shown capabilities in generating functional code, yet their utility in the development of clinical prediction tools has not been significantly explored. We evaluated GPT-4os capability to create a postoperative complication risk calculator similar to the existing National Surgical Quality Improvement Program (NSQIP) risk calculator. This included data preprocessing, predictive modeling, and development of a web application. Synthetic data of a similar structure to the NSQIP dataset was used when communicating with GPT-4o to maintain privacy. 512 lines of Python code were generated across 14 prompts, with one line requiring human editing. The resulting logistic regression models achieved similar Brier scores compared to the original NSQIP risk calculator and demonstrated strong discrimination (C-statistic > 0.75), while slightly underperforming previously reported predictive metrics for some outcomes. Development was completed in three hours. These findings suggest that LLMs can facilitate rapid development of clinical decision support tools, though output still requires human oversight and refinement.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 96%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 94%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 93%
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 93%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 93%
Similar papers in this journal
- LinkR: an open source, low-code and collaborative data science platform for healthcare data analysis and visualization 93%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 92%
- Predicting Prognosis in COVID-19 Patients using Machine Learning and Readily Available Clinical Data 91%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 93%
- Impact of the Federated Data Platform's digital surgery scheduling system on elective theatre utilisation at an NHS Trust: an interrupted time series analysis 92%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.