Back

AI-VOICE: A Method to Measure and Incorporate Patient Utilities Into AI-Informed Healthcare Workflows

Morse, K. E.; Higgins, M. C.; Qian, Y.; Callahan, A.; Shah, N. H.

2024-09-22 health informatics
10.1101/2024.09.19.24313990 medRxiv
Show abstract

BackgroundPatients are important participants in their medical care, yet artificial intelligence (AI) models are used to guide care with minimal patient input. This limitation is made partially worse due to a paucity of rigorous methods to measure and incorporate patient values of the tradeoffs inherent in AI applications. This paper presents AI-VOICE (Values-Oriented Implementation and Context Evaluation), a novel method to collect patient values, or utilities, of the downstream consequences stemming from an AI models use to guide care. The results are then used to select the models risk threshold, offering a mechanism by which an algorithm can concretely reflect patient values. MethodsThe entity being evaluated by AI-VOICE is an AI-informed workflow, which is composed of the patients health state, an action triggered by the AI model, and the benefits and harms accrued as a consequence of that action. The utilities of these workflows are measured through a survey-based, standard gamble experiment. These utilities define a patient-specific ratio of the cost of an inaccurate prediction versus the benefits of an accurate one. This ratio is mapped to the receiver-operator-characteristic curve to identify the risk threshold that reflects the patients values. The survey instrument is made freely available to researchers through a web-based application. ResultsA demonstration of AI-VOICE is provided using a hypothetical sepsis prediction algorithm. ConclusionAI-VOICE offers an accessible, quantitative method to incorporate patient values into AI-informed healthcare workflows.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.