Machine learning detects hidden treatment response patterns only in the presence of comprehensive clinical phenotyping
Auger, S. D.; Scott, G.
Show abstract
Inferential statistics traditionally used in clinical trials can miss relationships between clinical phenotypes and treatment responses. We simulated a randomised clinical trial to explore how gradient boosting (XGBoost) machine learning (ML) compares with traditional analysis when ground truth treatment responsiveness depends on the interaction of multiple phenotypic variables. As expected, traditional analysis detected a significant treatment benefit (outcome measure change from baseline = 4.23; 95% CI 3.64-4.82). However, recommending treatment based upon this evidence would lead to 56.3% of patients failing to respond. In contrast, ML correctly predicted treatment response in 97.8% (95% CI 96.6- 99.1) of patients, with model interrogation showing the critical phenotypic variables and the values determining treatment response had been identified. Importantly, when a single variable was omitted, accuracy dropped to 69.4% (95% CI 65.3-73.4). This proof of principle underscores the significant potential of ML to maximise the insights derived from clinical research studies. However, the effectiveness of ML in this context is highly dependent on the comprehensive capture of phenotypic data.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparing randomized trial designs to estimate treatment effect in rare diseases with longitudinal models: a simulation study showcased by Autosomal Recessive Cerebellar Ataxias using the SARA score 96%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 94%
- Comparison of Bayesian networks, G-estimation and linear models to estimate causal treatment effects in aggregated N-of-1 trials 93%
Similar papers in this journal
Similar papers in this journal
- A Scoping Review of Artificial Intelligence Applications in Clinical Trial Risk Assessment 95%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 94%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 94%
Similar papers in this journal
- Dynamic methods for ongoing assessment of site-level risk in risk-based monitoring of clinical trials: a scoping review 93%
- A modular pipeline for natural language processing-screened human abstraction of a pragmatic trial outcome from electronic health records 93%
- Estimating Counterfactual Placebo HIV Incidence in HIV Prevention Trials Without Placebo Arms Based on Markers of HIV Exposure 91%
Similar papers in this journal
- Machine learning approach to dynamic risk modeling of mortality in COVID-19: a UK Biobank study 94%
- Network and pathway expansion of genetic disease associations identifies successful drug targets 94%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.