High-fidelity discrimination of ARDS versus other causes of respiratory failure using natural language processing and iterative machine learning
Afshin-Pour, B.; Qiu, M.; Hosseini, S.; Stewart, M.; Horsky, J.; Aviv, R.; Zhang, N.; Narasimhan, M.; Chelico, J.; Musso, G.; Hajizadeh, N.
Show abstract
Despite the high morbidity and mortality associated with Acute Respiratory Distress Syndrome (ARDS), discrimination of ARDS from other causes of acute respiratory failure remains challenging, particularly in the first 24 hours of mechanical ventilation. Delay in ARDS identification prevents lung protective strategies from being initiated and delays clinical trial enrolment and quality improvement interventions. Medical records from 1,263 ICU-admitted, mechanically ventilated patients at Northwell Health were retrospectively examined by a clinical team who assigned each patient a diagnosis of "ARDS" or "non-ARDS" (e.g., pulmonary edema). We then applied an iterative pre-processing and machine learning framework to construct a model that would discriminate ARDS versus non-ARDS, and examined features informative in the patient classification process. Data made available to the model included patient demographics, laboratory test results from before the initiation of mechanical ventilation, and features extracted by natural language processing of radiology reports. The resulting model discriminated well between ARDS and non-ARDS causes of respiratory failure (AUC=0.85, 89% precision at 20% recall), and highlighted features unique among ARDS patients, and among and the subset of ARDS patients who would not recover. Importantly, models built using both clinical notes and laboratory test results out-performed models built using either data source alone, akin to the retrospective clinician-based diagnostic process. This work demonstrates the feasibility of using readily available EHR data to discriminate ARDS patients prospectively in a real-world setting at a critical time in their care and highlights novel patient characteristics indicative of ARDS.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Imputation of PaO2 from SpO2 values from the MIMIC-III Critical Care Database Using Machine-Learning Based Algorithms 97%
- An ML prediction model based on clinical parameters and automated CT scan features for COVID-19 patients 95%
- The incremental value of computed tomography of COVID-19 pneumonia in predicting ICU admission 95%
Similar papers in this journal
- Identification of physiological adverse events using continuous vital signs monitoring during paediatric critical care transport: a novel data-driven approach 95%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 94%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 93%
Similar papers in this journal
- Drivers of Mortality in COVID ARDS Depend on Patient Sub-Type 95%
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 94%
- Machine Learning Interpretability Methods to Characterize the Importance of Hematologic Biomarkers in Prognosticating Patients with Suspected Infection 93%
Similar papers in this journal
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 94%
- Machine learning in predicting respiratory failure in patients with COVID-19 pneumonia - challenges, strengths, and opportunities in a global health emergency 94%
- Multi-modal data to identify key factors influencing lung injury in ARDS patients undergoing invasive mechanical ventilation: A prospective observational study protocol 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.