AI-Driven Feature Selection Using Only Survey Variable Descriptions: Large Language Models Identify Adolescent Vaping Predictors
Zhang, K.; Zhao, Z.; Hu, Y.; Le, T.
Show abstract
ObjectiveTo evaluate the effectiveness of various Large Language Models (LLMs) in identifying reliable predictors of Electronic Nicotine Delivery Systems (ENDS) initiation among adolescents, using solely large-scale survey variable descriptions. MethodsA cohort of 7,943 tobacco-naive adolescents aged 12-16 years from the Population Assessment of Tobacco and Health (PATH) Study was analyzed to predict ENDS use at wave 5. Four instruction-tuned LLMs - GPT-4o, LLaMA 3.1-70B, Qwen 2.5-72B-Instruct, and DeepSeek-V3 - were systematically evaluated for text-based feature selection using only variable descriptions from wave 4.5. Selected features were used to train LightGBM classifiers, with model performance compared to a baseline. ResultsOur findings reveal notable consistency among the four instruction-tuned LLMs, with substantial overlap in the top predictors each model identified. These selected variables spanned critical domains such as peer and household influence, risk perception, and exposure to tobacco-related cues. LightGBM classifiers trained on PATH wave 4.5-5 data using features selected by the LLMs demonstrated strong predictive performance. Notably, Qwen 2.5-72B-Instruct achieved an AUC of 0.791 with 30 predictors, surpassing the baseline AUC of 0.768. DiscussionThe substantial overlap among the top predictors identified by different LLMs suggests a shared reasoning process, despite variations in model architecture and training. LightGBM classifiers trained on these LLM-selected features achieved performance comparable to, or exceeding, models trained on the full set of survey variables, underscoring the high quality of features selected solely from textual descriptions. Moreover, these findings are consistent with previous tobacco regulatory research, further validating the effectiveness of LLM-driven feature selection. ConclusionInstruction-tuned large language models can effectively perform text-based feature selection using survey variable descriptions alone, without accessing raw survey data. This scalable, interpretable, and privacy-preserving framework holds promise for behavioral health research and tobacco use surveillance.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Use of large language models as a scalable approach to understanding public health discourse 92%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 91%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 94%
- Causal Analysis for Multivariate Integrated Clinical and Environmental Exposures Data 93%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 92%
Similar papers in this journal
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 92%
- Empirical Sample Size Determination for Popular Classification Algorithms in Clinical Research 92%
- Optimal policy determination in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: A patient-physician digital twin dyad with deep Q-learning for treatment selection 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.