Back

Validation Of Natural Language Processing For Surgical Complication Surveillance: Detecting Eleven Postoperative Complications From Electronic Health Records

Dencker, E. E.; Bonde, A.; Troelsen, A.; Sillesen, M.

2025-04-07 surgery
10.1101/2025.04.07.25325367 medRxiv
Show abstract

BackgroundPostoperative complications (PCs) rates are crucial quality metrics in surgery, as they reflect both patient outcomes, perioperative care effectiveness and healthcare resource strain. Despite their importance, efficient, accurate, and affordable methods for tracking PCs are lacking. This study aimed to evaluate whether natural language processing (NLP) models could detect eleven PCs from surgical electronic health records (EHRs) at a level comparable to human curation. Methods17 486 surgical cases from 18 hospitals across two regions in Denmark, spanning six years, were included. The dataset was divided into training, validation, and test sets for NLP-model development and evaluation (50.2%/33.6%/16.2%). Model performance was compared against the current method of PC monitoring (ICD-10 codes) and manual curation, the latter serving as the gold standard. ResultsThe NLP-models had a ROC AUC between 0.901 to 0.999 for the test set. Sensitivity of the models when compared to manual curation ranged from 0.701 to 1.00, except for myocardial infarction (0.500). Positive Predictive Value (PPV) ranged from 0.0165 to 0.947, and Negative Predictive Value from 0.995 to 1.00. The NLP-models significantly outperformed ICD-10 coding in detecting PC, resulting in 16.3% of cases would require manual curation to reach a PPV of 1.00 ConclusionThe NLP models alone were able to detect PCs at an acceptable level and performed superior to ICD-10 codes. Combining NLP based and manual curation was required to reach a PPV of 1.00. Therefore, NLP algorithms present a potential solution for comprehensive and real-time monitoring of PCs across the surgical field.

Published in BMJ Connections Surgery · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.