Intricacies of Human-AI Interaction in Dynamic Decision-Making for Precision Oncology: A Case Study in Response-Adaptive Radiotherapy
Niraula, D.; Cuneo, K. C.; Dinov, I. D.; Gonzalez, B. D.; Jamaluddin, J. B.; Jin, J.; Luo, Y.; Matuszak, M. M.; Ten Haken, R. K.; Bryant, A. K.; Dilling, T. J.; Dykstra, M. P.; Frakes, J. M.; Liveringhouse, C. L.; Miller, S. R.; Mills, M. N.; Palm, R. F.; Regan, S. N.; Rishi, A.; Torres-Roca, J. F.; Yu, H.-H. M.; El Naqa, I.
Show abstract
BackgroundAdaptive treatment strategies that can dynamically react to individual cancer progression can provide effective personalized care. Longitudinal multi-omics information, paired with an artificially intelligent clinical decision support system (AI-CDSS) can assist clinicians in determining optimal therapeutic options and treatment adaptations. However, AI-CDSS is not perfectly accurate, as such, clinicians over/under reliance on AI may lead to unintended consequences, ultimately failing to develop optimal strategies. To investigate such collaborative decision-making process, we conducted a Human-AI interaction case study on response-adaptive radiotherapy (RT). MethodsWe designed and conducted a two-phase study for two disease sites and two treatment modalities--adaptive RT for non-small cell lung cancer (NSCLC) and adaptive stereotactic body RT for hepatocellular carcinoma (HCC)--in which clinicians were asked to consider mid-treatment modification of the dose per fraction for a number of retrospective cancer patients without AI-support (Unassisted Phase) and with AI-assistance (AI-assisted Phase). The AI-CDSS graphically presented trade-offs in tumor control and the likelihood of toxicity to organs at risk, provided an optimal recommendation, and associated model uncertainties. In addition, we asked for clinicians decision confidence level and trust level in individual AI recommendations and encouraged them to provide written remarks. We enrolled 13 evaluators (radiation oncology physicians and residents) from two medical institutions located in two different states, out of which, 4 evaluators volunteered in both NSCLC and HCC studies, resulting in a total of 17 completed evaluations (9 NSCLC, and 8 HCC). To limit the evaluation time to under an hour, we selected 8 treated patients for NSCLC and 9 for HCC, resulting in a total of 144 sets of evaluations (72 from NSCLC and 72 from HCC). Evaluation for each patient consisted of 8 required inputs and 2 optional remarks, resulting in up to a total of 1440 data points. ResultsAI-assistance did not homogeneously influence all experts and clinical decisions. From NSCLC cohort, 41 (57%) decisions and from HCC cohort, 34 (47%) decisions were adjusted after AI assistance. Two evaluations (12%) from the NSCLC cohort had zero decision adjustments, while the remaining 15 (88%) evaluations resulted in at least two decision adjustments. Decision adjustment level positively correlated with dissimilarity in decision-making with AI [NSCLC:{rho} = 0.53 (p < 0.001); HCC:{rho} = 0.60 (p < 0.001)] indicating that evaluators adjusted their decision closer towards AI recommendation. Agreement with AI-recommendation positively correlated with AI Trust Level [NSCLC:{rho} = 0.59 (p < 0.001); HCC:{rho} = 0.7 (p < 0.001)] indicating that evaluators followed AIs recommendation if they agreed with that recommendation. The correlation between decision confidence changes and decision adjustment level showed an opposite trend [NSCLC:{rho} = -0.24 (p = 0.045), HCC:{rho} = 0.28 (p = 0.017)] reflecting the difference in behavior due to underlying differences in disease type and treatment modality. Decision confidence positively correlated with the closeness of decisions to the standard of care (NSCLC: 2 Gy/fx; HCC: 10 Gy/fx) indicating that evaluators were generally more confident in prescribing dose fractionations more similar to those used in standard clinical practice. Inter-evaluator agreement increased with AI-assistance indicating that AI-assistance can decrease inter-physician variability. The majority of decisions were adjusted to achieve higher tumor control in NSCLC and lower normal tissue complications in HCC. Analysis of evaluators remarks indicated concerns for organs at risk and RT outcome estimates as important decision-making factors. ConclusionsHuman-AI interaction depends on the complex interrelationship between experts prior knowledge and preferences, patients state, disease site, treatment modality, model transparency, and AIs learned behavior and biases. The collaborative decision-making process can be summarized as follows: (i) some clinicians may not believe in an AI system, completely disregarding its recommendation, (ii) some clinicians may believe in the AI system but will critically analyze its recommendations on a case-by-case basis; (iii) when a clinician finds that the AI recommendation indicates the possibility for better outcomes they will adjust their decisions accordingly; and (iv) When a clinician finds that the AI recommendation indicate a worse possible outcome they will disregard it and seek their own alternative approach.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cluster-Based Toxicity Estimation of Osteoradionecrosis via Unsupervised Machine Learning: Moving Beyond Single Dose-Parameter Normal Tissue Complication Probability by Using Whole Dose-Volume Histograms for Cohort Risk Stratification 95%
- Bayesian Learning to Reduce Cardiac Risk for Locally Advanced NSCLC Patients Based on Personalized Radiotherapy Prescription 95%
- Initial Feasibility and Clinical Implementation of Daily MR-guided Adaptive Head and Neck Cancer Radiotherapy on a 1.5T MR-Linac System: Prospective R-IDEAL 2a/2b Systematic Clinical Evaluation of Technical Innovation 95%
Similar papers in this journal
- Artificial Intelligence Uncertainty Quantification in Radiotherapy Applications - A Scoping Review 95%
- First demonstration of the FLASH effect with ultrahigh dose-rate high-energy X-rays 94%
- Deep learning NTCP model for late dysphagia after radiotherapy for head and neck cancer patients based on 3D dose, CT and segmentations 94%
Similar papers in this journal
Similar papers in this journal
- Prediction of radiation-induced hypothyroidism using radiomic data analysis does not show superiority over standard normal tissue complication models 95%
- Standardising Breast Radiotherapy Structure Naming Conventions: A Machine Learning Approach 94%
- Effectiveness of FLASH vs conventional dose rate radiotherapy in a model of orthotopic, murine breast cancer 93%
Similar papers in this journal
- Large language models to help appeal denied radiotherapy services 95%
- Histology-based Prediction of Therapy Response to Neoadjuvant Chemotherapy for Esophageal and Esophagogastric Junction Adenocarcinomas Using Deep Learning 92%
- Early ctDNA kinetics as a dynamic biomarker of cancer treatment response 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.