Back

Development and Evaluation of an AI-Assisted, Privacy-Preserving Surgical Risk Calculator

Wolfrath, N.; SenthilKumar, G.; Ramamurthi, A.; Kothari, A. N.

2025-04-22 health informatics
10.1101/2025.04.18.25325614 medRxiv
Show abstract

Large language models (LLMs) have shown capabilities in generating functional code, yet their utility in the development of clinical prediction tools has not been significantly explored. We evaluated GPT-4os capability to create a postoperative complication risk calculator similar to the existing National Surgical Quality Improvement Program (NSQIP) risk calculator. This included data preprocessing, predictive modeling, and development of a web application. Synthetic data of a similar structure to the NSQIP dataset was used when communicating with GPT-4o to maintain privacy. 512 lines of Python code were generated across 14 prompts, with one line requiring human editing. The resulting logistic regression models achieved similar Brier scores compared to the original NSQIP risk calculator and demonstrated strong discrimination (C-statistic > 0.75), while slightly underperforming previously reported predictive metrics for some outcomes. Development was completed in three hours. These findings suggest that LLMs can facilitate rapid development of clinical decision support tools, though output still requires human oversight and refinement.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.