Real-world usage diminishes validity of Artificial Intelligence tools
Vaid, A.; Sawant, A.; Farinas, M. S.; Lee, J.; Kaul, S.; Kovatch, P.; Freeman, R.; Jiang, J.; Jayaraman, P.; Fayad, Z.; Argulian, E.; Lerakis, S.; Charney, A.; Wang, F.; Levin, M.; Glicksberg, B. S.; Narula, J.; Hofer, I.; Singh, K.; Nadkarni, G.
Show abstract
BackgroundSubstantial effort has been directed towards demonstrating use cases of Artificial Intelligence in healthcare, yet limited evidence exists about the long-term viability and consequences of machine learning model deployment. MethodsWe use data from 130,000 patients spread across two large hospital systems to create a simulation framework for emulating real-world deployment of machine learning models. We consider interactions resulting from models being re-trained to improve performance or correct degradation, model deployment with respect to future model development, and simultaneous deployment of multiple models. We simulate possible combinations of deployment conditions, degree of physician adherence to model predictions, and the effectiveness of these predictions. ResultsModel performance shows a severe decline following re-training even when overall model use and effectiveness is relatively low. Further, the deployment of any model erodes the validity of labels for outcomes linked on a pathophysiological basis, thereby resulting in loss of performance for future models. In either case, mitigations applied to offset loss of performance are not fully corrective. Finally, the randomness inherent to a system with multiple deployed models increases exponentially with adherence to model predictions. ConclusionsOur results indicate that model use precipitates interactions that damage the validity of deployed models, and of models developed in the future. Without mechanisms which track the implementation of model predictions, the true effect of model deployment on clinical care may be unmeasurable, and lead to patient data tainted by model use being permanently archived within the Electronic Healthcare Record.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 96%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 93%
- Community-acquired pneumonia identification from electronic health records in the absence of a gold standard: a Bayesian latent class analysis 92%
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 96%
- Developing Machine Learning Models for Predicting Intensive Care Unit Resource Use During the COVID-19 Pandemic 96%
- Emergency department admissions during COVID-19: explainable machine learning to characterise data drift and detect emergent health risks 96%
Similar papers in this journal
- Predicting hospital-onset COVID-19 infections using dynamic networks of patient contacts: an observational study 93%
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 93%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 92%
Similar papers in this journal
- sureLDA: A Multi-Disease Automated Phenotyping Method for the Electronic Health Record 94%
- High-throughput Phenotyping with Temporal Sequences 93%
- Personalizing renal replacement therapy initiation in the intensive care unit: a reinforcement learning-based strategy with external validation on the AKIKI randomized controlled trials 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.