Identifying Key Predictive Features for Opioid Use Disorder Using Machine Learning
Akhter, S.; Miller, J. H.
Show abstract
BackgroundOpioid Use Disorder (OUD) continues to pose a pressing public health challenge across the United States, highlighting the critical need for early and accurate risk assessment tools that facilitate prompt prevention and intervention efforts. Machine learning methods have emerged as valuable tools for parsing complex medical datasets and aiding in clinical decisions. However, their effectiveness and interpretability largely rely on the appropriateness and quality of selected input features. ObjectiveIn this work, we conducted a comprehensive comparison of three distinct feature selection strategies--Alternating Decision Tree (ADT)-based scoring, Cross-Validated Feature Evaluation (CVFE), and Hypergraph-Based Feature Evaluation (HFE)-- to identify the most predictive indicators of OUD. MethodsThe analysis was performed using data from the 2023 National Survey on Drug Use and Health (NSDUH), a dataset compiled by RTI International under the direction of the Substance Abuse and Mental Health Services Administration (SAMHSA). This dataset encompasses a broad spectrum of features related to demographics, behavior, mental health, and substance usage. Each feature selection method yielded a set of important predictors, which were subsequently used to train eXtreme Gradient Boosting (XGBoost) classification models. To enhance model transparency and interpretability, SHapley Additive exPlanations (SHAP) was employed to illustrate the influence of individual variables on model predictions. ResultsThe performance of the models was evaluated and compared, with the model informed by CVFE-selected features achieving the best outcomes--demonstrating a predictive accuracy of 79.11% and an area under the curve (AUC) of 0.8652. The top 10 most influential features, based on SHAP value rankings from the best-performing model, included past-year misuse of pain relievers, recent alcohol use disorder, age group, history of asthma, receipt of substance use treatment in the past year, educational attainment, household size, total household income, marital status, and race/ethnicity. The web application, accessible via https://shiny.tricities.wsu.edu/oud-prediction/, offers prediction outcomes, probability metrics, and a SHAP visualization generated from the best model built using cross-validation-based approach. ConclusionsThe findings highlight the crucial importance of effective feature selection in enhancing both model accuracy and interpretability, ultimately supporting the development of practical, data-driven approaches that may help healthcare providers assess OUD risk and tailor prevention strategies to individual needs. Trial registrationNot applicable as this research is not a clinical trial.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 95%
- Assessing the effects of data drift on the performance of machine learning models used in clinical sepsis prediction 92%
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests 92%
Similar papers in this journal
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 93%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 92%
- Framework for Identifying Drug Repurposing Candidates from Observational Healthcare Data 92%
Similar papers in this journal
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 94%
- Explainable AI enables clinical trial patient selection to retrospectively improve treatment effects in schizophrenia 93%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 93%
Similar papers in this journal
- Empirical Sample Size Determination for Popular Classification Algorithms in Clinical Research 95%
- Optimal policy determination in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: A patient-physician digital twin dyad with deep Q-learning for treatment selection 92%
- An Interpretable Machine Learning Framework for Accurate Severe vs Non-severe COVID-19 Clinical Type Classification 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.