Auditor Models to Suppress Poor AI Predictions Can Improve Human-AI Collaborative Performance
Brown, K.; Wrenn, J. O.; Jackson, N. J.; Cauley, M. R.; Collins, B.; Novak, L. L.; Malin, B. A.; Ancker, J. S.
Show abstract
ObjectiveHealthcare decisions are increasingly made with the assistance of machine learning (ML). ML has been known to have unfairness - inconsistent outcomes across subpopulations. Clinicians interacting with these systems can perpetuate such unfairness by overreliance. Recent work exploring ML suppression - silencing predictions based on auditing the ML - shows promise in mitigating performance issues originating from overreliance. This study aims to evaluate the impact of suppression on collaboration fairness and evaluate ML uncertainty as desiderata to audit the ML. Materials and MethodsWe used data from the Vanderbilt University Medical Center electronic health record (n = 58,817) and the MIMIC-IV-ED dataset (n = 363,145) to predict likelihood of death or ICU transfer and likelihood of 30-day readmission. Our simulation study used gradient-boosted trees as well as an artificially high-performing oracle model. We derived clinician decisions directly from the dataset and simulated clinician acceptance of ML predictions based on previous empirical work on acceptance of CDS alerts. We measured performance as area under the receiver operating characteristic curve and algorithmic fairness using absolute averaged odds difference. ResultsWhen the ML outperforms humans, suppression outperforms the human alone (p < 0.034) and at least does not degrade fairness. When the human outperforms the ML, suppression outperforms the human (p < 5.2 x 10-5) but the human is fairer than suppression (p < 0.0019). Finally, incorporating uncertainty quantification into suppression approaches can improve performance. ConclusionSuppression of poor-quality ML predictions through an auditor model shows promise in improving collaborative human-AI performance and fairness.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 95%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 95%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 95%
Similar papers in this journal
Similar papers in this journal
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 94%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.