Back

Physician gestalt compared with AI model to predict intubation in critically ill patients

Miller, M. A.; Lu, X.; Pearce, A. K.; Malhotra, A.; Nemati, S. A.

2025-12-21 intensive care and critical care medicine
10.64898/2025.12.19.25342667 medRxiv
Show abstract

RationaleIntubation and mechanical ventilation are associated with high mortality. Accurately predicting which patients are at the highest risk of intubation can enable interventions to reduce their risk. The performance of intensive care physicians to predict the need for intubation within the next 24 hours for medically critically ill patients is unknown. Machine learning models are adept at prediction tasks. ObjectiveIn this study, we perform a prospective observational study in two ICUs to survey intensivists to test their accuracy at predicting the need for intubation within 24 hours of patients under their care. Physician predictions of intubation are then compared to predictions from a machine learning model called Vent.io. MethodsPrimary metrics included prediction sensitivity, specificity, and descriptive statistics for both physician and machine learning model. Generalized linear mixed models were developed to investigate the fixed effect of the predictor (physician vs Vent.io) on both sensitivity and specificity while accounting for the random effects from different physicians and reported by odds ratio and 95% confidence interval. Similar modeling was also used to test the relationship between physician confidence and correctness. ResultsOverall, physicians are quite confident in their predictions of intubation with a median score of 8 (on a 0-10 point scale, with 0 being not at all confident and 10 being extremely confident) out of the 302 surveys administered. Sensitivity was 0.190 and 0.714 for physicians and Vent.io, respectively. Specificity was 0.960 and 0.673 for physicians and Vent.io, respectively. Generalized linear mixed modeling showed that physician confidence is associated with greater odds of correctly predicting intubation outcome (OR 1.49; 95% CI 1.22-1.84; p<.001). Vent.io had significantly greater odds of being correct when patients required intubation compared to physicians (OR 18.68; 95% CI 1.87-186.31; p=0.013). However, intensive care physicians outperformed Vent.io at correctly predicting when patients did not require intubation (OR 24.80; 95% CI 13.22-46.52; p<0.001). ConclusionsWhile promising, Vent.io needs real-time testing in a randomized clinical trial to determine if its deployment can improve clinical outcomes.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.