Methodological Guidance for Predictor Variable Selection for Adolescent Smoking Outcomes in Global Youth Tobacco Survey Using R and Python
Ng'ambi, W. F.; Zyambo, C.; Kazembe, L.
Show abstract
BackgroundThe Global Youth Tobacco Survey (GYTS) is widely used to monitor tobacco use among adolescents worldwide. However, inconsistent analytical approaches particularly in handling complex survey designs and predictor selection limit comparability across countries, survey waves, and software platforms. Although much of the GYTS literature relies on proprietary tools such as SAS and SPSS, practical and transparent guidance on implementing reproducible, theory-informed analyses remains limited. A unified workflow that respects the surveys design while supporting cross-platform implementation is needed. MethodsWe developed a reproducible, open-source workflow for analysing GYTS data using R and Python. In R, analyses were conducted using the survey package (svydesign and svyglm) with constrained stepwise selection via stepAIC. In Python, a custom constrained stepwise procedure was implemented using statsmodels generalized linear models. The workflow explicitly incorporates survey weights, stratification, and clustering; harmonises variables across countries; protects a priori demographic covariates; and ensures consistent treatment of categorical predictors. The approach is illustrated using data from Zambia (n = 2,959) and pooled data from Ghana, Mauritius, Seychelles, and Togo (n = 15,914). Predictor selection was guided by Social Cognitive Theory and evidence from systematic reviews. ResultsThe constrained selection framework consistently retained key demographic variables (age, sex, and grade) while allowing data-driven selection of modifiable predictors using the Akaike Information Criterion. When identical constraints were applied, the R and Python implementations selected identical models and produced nearly equivalent point estimates (adjusted odds ratio differences <0.01), although Python-based confidence intervals did not account for clustering. Of 18 candidate predictors across individual, social, media, and policy domains, 14 were retained. The strongest independent predictors included awareness of tobacco products (OR = 5.61, 95% CI: 4.65- 6.78), peer smoking (OR = 4.57, 95% CI: 3.34-6.25), and exposure to tobacco marketing (OR = 2.34, 95% CI: 1.89-2.91). ConclusionsThis study provides a generalisable, theory-informed framework for predictor selection in complex survey data using open-source tools. The workflow supports consistent analyses across countries, survey waves, and software platforms, and is transferable to other youth and adult population surveys. All code and harmonisation resources are openly available to support reproducibility and adaptation. Plain-Language SummaryO_LIWhat we asked: Can we predict adolescent smoking using GYTS data in a way that is easy to follow and reproducible across software? C_LIO_LIWhat we did: Built a single workflow that respects survey design (weights, strata, clusters) and selects predictors using four explicit criteria: theoretical grounding in Social Cognitive Theory, empirical support from prior studies, relevance for intervention, and cross-country validity. Core demographics (age, sex, grade, region) were protected as essential confounders, while other predictors were selected based on statistical fit. The workflow runs equivalently in R and Python. C_LIO_LIWhy it matters: Many GYTS studies use weights only and ignore clustering and stratification, which makes confidence intervals too narrow. More importantly, most analyses include variables arbitrarily or let software drop important confounders automatically. Our approach ensures theoretically meaningful, policy-relevant variables are retained, producing more reliable and actionable results for prevention programs. C_LI
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing Compliance with Smoke-Free Laws in Purbakhola Rural Municipality, Nepal: A Cross-Sectional Observational Study 91%
- Intimate partner violence and its correlates in middle-aged and older adults during the COVID-19 pandemic: A multi-country secondary analysis 91%
- Mechanisms for the prevention of adolescent intimate partner violence: a realist review of interventions in low- and middle-income countries 90%
Similar papers in this journal
- A factorial randomised controlled trial to examine the potential effect of a text-message based intervention on reducing adolescent susceptibility to e-cigarette use: A study protocol 93%
- Economic and social impacts of COVID-19 and public health measures: results from an anonymous online survey in Thailand, Malaysia, the United Kingdom, Italy and Slovenia 92%
- Describing the inputs, activities and outputs of "10,000 Lives", a coordinated regional smoking cessation initiative in Central Queensland, Australia 91%
Similar papers in this journal
- Youth susceptibility to tobacco use: Is it general or specific? 92%
- Characterization of Trajectories of Physical Activity and Cigarette Smoking from Early Adolescence to Adulthood 91%
- How can we maximise the benefits of smoke-free prisons? Decision analytic model to predict potential impacts on public health 91%
Similar papers in this journal
- Transitions from smoking to exclusive e-cigarette use, dual use, or stopping nicotine use in ALSPAC and their association with modifiable and sociodemographic factors 93%
- Changes in vaping trends since the announcement of an impending ban on disposable vapes: a population study in Great Britain 93%
- Moderators of changes in smoking, drinking, and quitting behaviour associated with the first Covid-19 lockdown in England 93%
Similar papers in this journal
- Effectiveness of Online Training in Improving Primary Care Doctors Competency in Brief Tobacco Interventions: A Cluster Randomised Controlled Trial of WHO Modules in Delta State, Nigeria 92%
- Investigating the added value of biomarkers compared with self-reported smoking in predicting future e-cigarette use: Evidence from a longitudinal UK cohort study 92%
- Variation in English Covid booster uptake 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.