Classification of Adolescent Drinking via Behavioral, Biological, and Environmental Features: A Machine Learning Approach with Bias Control
Liu, R.; Azzam, M.; Zabik, N.; Wan, S.; Blackford, J.; Wang, J.
Show abstract
In 2024, approximately 30% of U.S. adolescents reported having consumed alcohol at least once in their lifetime, with about 25% of these individuals engaging in binge drinking. Adolescent alcohol use is associated with neurodevelopmental impairments, elevated risk of later alcohol use, and mental health disorders. These findings underscore the importance of identifying the variables driving adolescent alcohol use and leveraging them for early identification and targeted intervention. Previous studies have typically developed machine-learning classification models that use neuroimaging data in combination with limited clinical measurements. Neuroimaging data are expensive and difficult to obtain at scale, whereas clinical measures are more practical for large-scale screening due to their low cost and widespread accessibility. However, clinical-only approaches for alcohol drinking classification remain largely underexplored. Furthermore, prior studies have often focused on adults, limiting generalizability to the broader adolescent population. Additionally, confounding factors such as age and substance use, which are strongly correlated with alcohol consumption, have often been inadequately addressed, potentially inflating classification performance. Finally, class imbalance remains a persistent challenge, with prior attempts yielding only limited improvements. To address these limitations, we propose FocalTab, a framework that integrates TabPFN with focal loss for robust generalization and effective mitigation of class imbalance. The approach also incorporates an initial preprocessing step to remove confounding factors to account for age and substance-use. We compare FocalTab against state-of-the-art methods across different variable selections and dataset settings. FocalTab achieves the highest accuracy (84.3%) and specificity (80.0%) in the most stringent setting, in which both age and substance use variables were excluded, whereas competing models drop to near-chance specificity (12-24%). We further applied SHapley Additive exPlanations (SHAP) analysis to identify key clinical predictors of drinker classification, supporting enhanced screening and early intervention.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A Multivariate Approach to Understanding the Genetic Overlap between Externalizing Phenotypes and Substance Use Disorders 90%
- Cue-induced effects on decision-making distinguish subjects with gambling disorder from healthy controls 89%
- Forward planning in a population-based alcohol use disorder sample 89%
Similar papers in this journal
- Specific diagnostic criteria identify those at high risk for progression from ‘preaddiction’ to severe alcohol use disorder 89%
- Effectiveness of a Text Message Intervention Promoting Seat Belt Use Among Targeted Young Adults: A Randomized Clinical Trial 89%
- Neuroanatomical variability associated with early substance use initiation: Results from the ABCD Study 88%
Similar papers in this journal
- Modelling Alcohol Consumption Patterns to Enable Policy Impact Assessment 92%
- Can alcohol consumption in Germany be reduced by alcohol screening, brief intervention and referral to treatment in primary health care? Results of a simulation study 92%
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 92%
Similar papers in this journal
- Affinity Scores: An Individual-centric Fingerprinting Framework for Neuropsychiatric Disorders 91%
- Computational Mechanisms Underlying Multi-Step Planning Deficits in Methamphetamine Use Disorder 90%
- DNA Methylation as a Potential Mediator of the Association Between Prenatal Tobacco and Alcohol Exposure and Child Neurodevelopment in a South African Birth Cohort 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.