Can accurate demographic information about people who use prescription medications non-medically be derived from Twitter big data?
Yang, Y.-C.; Al-Garadi, M. A.; Love, J. S.; Cooper, H. L. F.; Perrone, J.; Sarker, A.
Show abstract
Traditional surveillance mechanisms for nonmedical prescription medication use (NPMU) involve substantial lags. Social media-based approaches have been proposed for conducting close-to-real-time surveillance, but such methods typically cannot provide fine-grained statistics about subpopulations. We address this gap by developing methods for automatically characterizing a large Twitter NPMU cohort (n=288,562) in terms of age-group, race, and gender. Our methods achieved 0.88 precision (95%-CI: 0.84-0.92) for age-group, 0.90 (95%-CI: 0.85-0.95) for race, and 0.94 accuracy (95%-CI: 0.92-0.97) for gender. We compared the automatically-derived statistics for the NPMU of tranquilizers, stimulants, and opioids from Twitter to statistics reported in traditional sources (eg., the National Survey on Drug Use and Health). Our estimates were mostly consistent with the traditional sources, except for age-group-related statistics, likely caused by differences in reporting tendencies and representations in the population. Our study demonstrates that subpopulation-specific estimates about NPMU may be automatically derived from Twitter to obtain early insights.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automatic Gender Detection in Twitter Profiles for Health-related Cohort Studies 98%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 93%
- Characterization and Racial Stratification of Social Determinants of Health for Individuals with Type 2 Diabetes as Recorded in Electronic Health Records: Implications for Artificial Intelligence Development 92%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 93%
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests 91%
- Emergence and Evolution of Big Data Analytics in HIV Research: Bibliometric Analysis of Federally Sponsored Studies 2000-2019 90%
Similar papers in this journal
- Which social media platforms facilitate monitoring the opioid crisis? 95%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 91%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.