AI is Smart. Is it Wise? Quantifying the Effect of Patient-Choice (β) on Physical Outcomes
Gurel, O.; Rasmussen, M. F.; Veginati, V.; Weinstein, J. N.
Show abstract
Large language models (LLMs) increasingly guide clinical decisions through population-level evidence, yet they cannot encode individual patient preferences. When treatments yield comparable outcomes, patient choice may drive decisions, though its effect remains unquantified. The Spine Patient Outcomes Research Trial (SPORT)--marked by similar surgical and nonoperative results and substantial crossover--provided a natural experiment to use causal-inference methods to estimate unbiased treatment effects and quantify the contribution of patient choice to outcomes. Using only published aggregate results from SPORT, we conducted two-stage least squares instrumental-variable analysis using randomized treatment assignment as the instrument, with Complier Average Causal Effects (CACE) and E-values assessing sensitivity to unmeasured confounding. Primary outcomes were SF-36 Bodily Pain, SF-36 Physical Function scores, and the Oswestry Disability Index. We decomposed treatment effects into , the biological treatment mechanism, and {beta}, the patient-choice contribution. Aggregate estimates revealed G = 15.7 (0.5) and {beta}G = 7.4 (3.4), with the net difference between surgical and nonoperative treatment effects {Delta} {approx} 0.65. This analysis quantifies a measurable and significant effect of patient choice ({beta}) on physical outcomes. When treatment effects are comparable ({Delta} small), {beta}--a dimension inaccessible to current LLMs trained on -biased population-level evidence--emerges as the dominant driver of decision-making. These findings provide an empirical grounding for informed choice, clarify the limits of LLMs trained on -biased evidence, and quantify a structural constraint in AI-driven clinical decision support. Key messagesO_LIThe effect of patient choice ({beta}) on physical outcomes is real, measurable, and clinically meaningful. C_LIO_LI{beta} becomes the dominant driver of outcomes when biological treatment differences ({Delta}) are small. C_LIO_LILLMs cannot encode {beta} because they are trained on -biased population-level evidence. C_LIO_LIThese findings provide the empirical foundation for informed choice--not just informed consent. C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Modeling trajectories of routine blood tests as dynamic biomarkers for outcome in spinal cord injury 91%
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 90%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 90%
Similar papers in this journal
- Evidence of Unreliable Data and Poor Data Provenance in Clinical Prediction Model Research and Clinical Practice 90%
- Tool to assess risk of bias due to missing evidence in network meta-analysis (ROB-MEN): elaboration and examples 89%
- Evidence of unexplained discrepancies between planned and conducted statistical analyses: a review of randomized trials 89%
Similar papers in this journal
- Risk factors for heart failure with preserved or reduced ejection fraction among Medicare beneficiaries: Applications of competing risks analysis and gradient boosted model. 88%
- The long-term impact of vaginal surgical mesh devices in UK primary care: a cohort study in the CPRD 87%
- Is age the most important risk factor in COVID-19 patients? The relevance of comorbidity burden: A retrospective analysis of 10,090 hospitalizations 87%
Similar papers in this journal
- Large-scale validation of the Prediction model Risk Of Bias ASsessment Tool (PROBAST) using a short form: high risk of bias models show poorer discrimination 92%
- Estimating and Testing an Index of Bias Attributable to Composite Outcomes in Comparative Studies 91%
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 90%
Similar papers in this journal
- Unequal Recovery in Colorectal Cancer Screening Following the COVID-19 Pandemic: A Comparative Microsimulation Analysis 90%
- Machine learning analysis of a digital insole versus clinical standard gait assessments for digital endpoint development 90%
- The Proportion of Randomized Controlled Trials That Inform Clinical Practice: A Longitudinal Cohort Study of Trials Registered on ClinicalTrials.gov 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.