Back

Large Language Model Embeddings of Surgical Procedural Names for Confounding Adjustment in Perioperative Observational Studies

Han, L.; Ghanem, M.; Simhambhatla, M. K.; Chung, P.; Aghaeepour, N.

2026-08-25 anesthesia
10.64898/2026.08.22.26361065 medRxiv
Show abstract

Background: Perioperative observational studies are increasingly used to evaluate anesthesia practices that are difficult to test in randomized trials, but treatment selection is influenced by surgical procedure type. Scalable and robust methods are needed to adjust for procedure-level confounding across heterogeneous surgical cohorts. Methods: We developed and assessed the utility of a surgical-name embedding framework using free-text procedure names from 627,624 adult perioperative records. Procedure names were embedded using open-source sentence-embedding models, reduced with principal components analysis, and incorporated into entropy-balanced observational analyses. We compared unweighted, clinical covariate adjustment, and clinical covariate plus surgical-name embedding adjustment across three replications of recent perioperative randomized trials: GA-CARES for total intravenous versus volatile anesthesia on cancer mortality, PADDI for dexamethasone and surgical site infection, and GAP for perioperative gabapentin and postoperative length of stay. Results: In the GA-CARES replication, unweighted and clinical covariate adjustment suggested lower two-year mortality with total intravenous anesthesia, whereas adding surgical-name embeddings attenuated the estimate to a nonsignificant association consistent with the randomized trial (OR 0.84 [95% CI, 0.62-1.14]; p=0.274). In the PADDI replication, adding surgical-name adjustment reproduced the trial's overall non-harm finding for 30-day surgical site infection (OR 0.72 [95% CI 0.69-0.76], p<0.001) while preserving the expected protective association with postoperative nausea and/or vomiting. In the GAP replication, the full cohort was null across adjustment strategies, but surgical-name embeddings were required to recover trial-consistent null length-of-stay estimates across cardiac, thoracic, and abdominal subgroups. Department indicators and randomly generated covariates did not reproduce the effect of surgical-name embeddings, supporting the presence and importance of procedure-specific information. Conclusions: Free-text surgical-name embeddings provide a scalable method for representing surgical context in perioperative observational studies. Surgical-name adjustment improved concordance with randomized trial benchmarks while preserving expected treatment effects, supporting its use as an additional layer of confounding adjustment in large perioperative datasets.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.