Employing Data-Driven Techniques to Explore the Lay Public's Health Concerns with Vaping E-Cigarettes
Ayers, J. W.; Poliak, A.; DeLucia, A.; Zhu, Z.; Pitts, S.; Navarro, M.; Shojaie, S.; Dredze, M.
Show abstract
While the public perceives e-cigarettes as less harmful than combustible tobacco, little is known about their specific health concerns regarding vaping. We demonstrate a data-driven strategy to discover the public's health concerns about vaping e-cigarettes expressed on social media. We obtained all public posts from the largest e-cigarette-related subreddit, r/electronic_cigarette, from its inception on September 17, 2008, through April 1, 2022 (N = 10,403,433). We identified health concerns attributed to vaping by (a) selecting all cause phrases containing "cause" and its inflections, (b) calculating the empirical frequency ratio of words and bi-grams occurring in these phrases relative to random phrases, (c) retaining the 10% of words with the greatest empirical frequency of occurring in cause phrases, and (d) annotating this sample for health-relevant concerns and their subjects. In total, 76,342 posts contained cause phrases, with increased volume over time. Of the 425 words most strongly associated with cause phrases compared to random phrases, 53.4% (95%CI, 48.7-58.1) were identified as health-relevant. The top health-related concern was lipoid pneumonia, cited in 5.9% (95%CI, 5.0-6.8) of all cause phrases, followed by pneumonia (4.1%; 95%CI, 3.3-4.9), and nausea (2.7%;95%CI, 2.0-3.4). The top health concern subjects were respiratory, representing 23.7% (95%CI, 18.5-29.5) of all cause phrases, followed by gastrointestinal (12.7%; 95%CI, 8.8-17.2) and cardiovascular (8.5%; 95%CI, 5.3-12.3) concerns. Other subjects included neurological, dermatological, oral health, sexual health, psychiatric, oncologic, addiction, and sleep concerns. Because our strategy relies on data-driven techniques, our analysis can be integrated into routine social media monitoring and applied across different types of social media and text data, potentially leading to more timely identification of emerging concerns and a broader understanding across platforms. As a result, experts can craft messaging that accounts for current perceptions held by the public using our method.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 94%
- Tracking private WhatsApp discourse about COVID-19: A longitudinal infodemiology study in Singapore 93%
- Global Infodemiology of COVID-19: Focus on Google web searches and Instagram hashtags 92%
Similar papers in this journal
- Computational network models for forecasting and control of mental health trajectories in digital applications 90%
- Digital Health Tools for the Passive Monitoring of Depression: A Systematic Review of Methods 90%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 89%
Similar papers in this journal
Similar papers in this journal
- Utility and limitations of Google searches for tracking disease: the case of taste and smell loss as markers for COVID-19 91%
- Sociodemographic Characteristics of Missing Data in Digital Phenotyping 90%
- Mapping internet activity in Australian cities during COVID-19 lockdown: how occupational factors drive inequality 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.