Back

Journal of Medical Internet Research

JMIR Publications Inc.

All preprints, ranked by how well they match Journal of Medical Internet Research's content profile, based on 87 papers previously published here. The average preprint has a 0.11% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Dynamic Topic Alignment and Sentiment between Official Health Communication and General Public Discourse during COVID-19: A Comprehensive Infoveillance Framework

Yin, S.; Xin, W.; Chen, S.; Ge, Y.

2026-05-27 public and global health 10.64898/2026.05.23.26353966 medRxiv
Top 0.1%
67.1%
Show abstract

Social media has become a critical channel for public health communication during the COVID-19 pandemic, yet how official health messaging aligns with broader public discourse remains insufficiently understood. This study develops an end-to-end info-veillance framework to examine the dynamic relationship between Centers for Disease Control and Prevention (CDC) communications and general public discourse on social media. We analyzed 17,524 CDC tweets and 67,895 public discourse tweets. Biterm Topic Model (BTM) was used to extract topics from each corpus, and a novel topic consistency scoring system integrating cosine similarity with daily public topic prominence was developed to quantify temporal alignment between official health communication and public discourse. Two complementary sentiment measures were incorporated: expected sentiment (average emotional tone) and net sentiment (overall emotional intensity). Temporal relationships were examined using autoregressive integrated moving average with exogenous variables (ARIMAX) models. Results show that topic alignment increased over time across CDC topics, while expected sentiment remained consistently negative. Higher alignment was associated with immediate and delayed changes in expected sentiment and stronger emotional intensity in net sentiment based on ARIMAX results. These findings suggest that topic alignment reflects public attention rather than agreement with official communications, and is associated with more negative emotional responses. This framework provides a scalable, generalizable approach to investigate and evaluate public engagement with official health communication.

2
Infoveillance study on the dynamic associations between CDC social media contents and epidemic measures during COVID-19

Yin, S.; Chen, S.; Ge, Y.

2023-06-27 health informatics 10.1101/2023.06.26.23291921 medRxiv
Top 0.1%
60.1%
Show abstract

BackgroundHealth agencies have been widely adopting social media to disseminate important information, educate the public on emerging health issues, and understand public opinions. The Centers for Disease Control and Prevention (CDC) has been one of the leading agencies that utilizes social media platforms during the COVID-19 pandemic to communicate with the public and mitigate the disease in the United States. It is crucial to understand the relationships between CDCs social media communication and the actual epidemic metrics to improve public health agencies communication strategies during health emergencies. ObjectiveThe aim of this study was to identify key topics in tweets posted by CDC during the pandemic, to investigate the temporal dynamics between these key topics and the actual COVID-19 epidemic measures, and to make recommendations for CDCs digital health communication strategies for future health emergencies. MethodsTwo types of data were collected: 1) a total of 17,524 COVID-19-related English tweets posted by the CDC between December 7, 2019 and January 15, 2022; 2) COVID-19 epidemic measures in the U.S. from the public GitHub repository of Johns Hopkins University from January 2020 to July 2022. Latent Dirichlet allocation (LDA) topic modeling was applied to identify key topics from all COVID-19-related tweets posted by CDC, and the final topics were determined by domain experts. Various multivariate time series analysis techniques were applied between each of the identified key topics and actual COVID-19 epidemic measures to quantify the dynamic associations between these two types of time series data. ResultsFour major topics from CDCs COVID-19 tweets were identified: 1) information on prevention of health outcomes of COVID-19; 2) pediatric intervention and family safety; 3) updates of the epidemic situation of COVID-19; 4) research and community engagement to curb COVID-19. Multivariate analyses showed that there were significant variabilities of progression between CDCs topics and the actual COVID-19 epidemic measures. Some CDCs topics showed substantial associations with the COVID-19 measures over different time spans throughout the pandemic, expressing similar temporal dynamics between these two types of time series data. ConclusionsOur study is the first to comprehensively investigate the dynamic associations between topics discussed by CDC on Twitter and the COVID-19 epidemic measures in the U.S. We identified four major topic themes via topic modeling and explored how each of these topics was associated with each major epidemic measure by performing various multivariate time series analyses. We recommend that it is critical for public health agencies, such as CDC, to disseminate and update timely and accurate information to the public and align major topics with the key epidemic measures over time. We suggest that social media can help public health agencies to inform the public on health emergencies and to mitigate them effectively.

3
Mining Twitter to Assess the Determinants of Health Behavior towards Palliative Care in the United States

Zhao, Y.; Zhang, H.; Huo, J.; Guo, Y.; Wu, Y.; Bian, J.

2020-03-30 health informatics 10.1101/2020.03.26.20038372 medRxiv
Top 0.1%
49.9%
Show abstract

Palliative care is a specialized service with proven efficacy in improving patients quality-of-life. Nevertheless, lack of awareness and misunderstanding limits its adoption. Research is urgently needed to understand the determinants (e.g., knowledge) related to its adoption. Traditionally, these determinants are measured with questionnaires. In this study, we explored Twitter to reveal these determinants guided by the Integrated Behavioral Model. A secondary goal is to assess the feasibility of extracting user demographics from Twitter data--a significant shortcoming in existing studies that limits our ability to explore more fine-grained research questions (e.g., gender difference). Thus, we collected, preprocessed, and geocoded palliative care-related tweets from 2013 to 2019 and then built classifiers to:1) categorize tweets into promotional vs. consumer discussions, and 2) extract user gender. Using topic modeling, we explored whether the topics learned from tweets are comparable to responses of palliative care-related questions in the Health Information National Trends Survey.

4
Social Media Insights Into US Mental Health Amid the COVID-19 Pandemic. A Longitudinal Twitter Analysis (JANUARY-APRIL 2020)

Valdez, D.; ten Thij, M.; Bathina, K. C.; Rutter, L. A.; Bollen, J.

2020-12-10 public and global health 10.1101/2020.12.01.20241943 medRxiv
Top 0.1%
45.9%
Show abstract

BackgroundThe COVID-19 pandemic led to unprecedented mitigation efforts that disrupted the daily lives of millions. Beyond the general health repercussions of the pandemic itself, these measures also present a significant challenge to the worlds mental health and healthcare systems. Considering traditional survey methods are time-consuming and expensive, we need timely and proactive data sources to respond to the rapidly evolving effects of health policy on our populations mental health. Significant pluralities of the US population now use social media platforms, such as Twitter, to express the most minute details of their daily lives and social relations. This behavior is expected to increase during the COVID-19 pandemic, rendering social media data a rich field from which to understand personal wellbeing. PurposeBroadly, this study answers three research questions: RQ1: What themes emerge from a corpus of US tweets about COVID-19?; RQ2: To what extent does social media use increase during the onset of the COVID-19 pandemic?; and RQ3: Does sentiment change in response to the COVID-19 pandemic? MethodsWe analyzed 86,581,237 public domain English-language US tweets collected from an open-access public repository in three steps1. First, we characterized the evolution of hashtags over time using Latent Dirichlet Allocation (LDA) topic modeling. Second, we increased the granularity of this analysis by downloading Twitter timelines of a large cohort of individuals (n = 354,738) in 20 major US cities to assess changes in social media use. Finally, using this timeline data, we examined collective shifts in public mood in relation to evolving pandemic news cycles by analyzing the average daily sentiment of all timeline tweets with the Valence Aware Dictionary and sEntiment Reasoner (VADER) sentiment tool2. ResultsLDA topics generated in the early months of the dataset corresponded to major COVID-19 specific events. However, as state and municipal governments began issuing stay-at-home orders, latent themes shifted towards US-related lifestyle changes rather than global pandemic-related events. Social media volume also increased significantly, peaking during stay-at-home mandates. Finally, VADER sentiment analysis sentiment scores of user timelines were initially high and stable, but decreased significantly, and continuously, by late March. Discussion & ConclusionOur findings underscore the negative effects of the pandemic on overall population sentiment. Increased usage rates suggest that, for some, social media may be a coping mechanism to combat feelings of isolation related to long-term social distancing. However, in light of the documented negative effect of heavy social media usage on mental health, for many social media may further exacerbate negative feelings in the long-term. Thus, considering the overburdened US mental healthcare structure, these findings have important implications for ongoing mitigation efforts.

5
Understanding Public Attitudes Towards Human Papillomavirus Vaccination in Japan: Insights from Social Media Stance Analysis Using Large Language Models

Niu, Q.; Liu, J.

2024-10-07 health informatics 10.1101/2024.10.07.24315018 medRxiv
Top 0.1%
45.8%
Show abstract

BackgroundDespite the reinstatement of proactive human papillomavirus (HPV) vaccine recommendations in 2022, Japan continues to face persistently low HPV vaccination rates, posing significant public health challenges. Misinformation, complacency, and accessibility issues have been identified as key factors undermining vaccine uptake. ObjectiveThis study aims to understand how factors such as misinformation, public health events, and attitudes toward other vaccines, like COVID-19, influence HPV vaccine hesitancy, by analyzing the evolution of public attitudes towards HPV vaccination in Japan by examining social media content. MethodsWe collected tweets related to HPV vaccine from 2011 to 2021. Traditional natural language processing (NLP) methods and large language models (LLMs) was utilized to perform stance analysis on collected data. The analysis included stance identification, time series analysis, topic modeling, and logic analysis. We framed our findings within the context of the WHOs 3Cs model--Confidence, Complacency, and Convenience. ResultsPublic confidence in the HPV vaccine fluctuated in response to government policies and media events, with misinformation playing a critical role in eroding trust. Complacency increased following the suspension of recommendations in 2013 but decreased as advocacy resumed in 2020. Accessibility (Convenience) was also found to be a key determinant of vaccination uptake. HPV vaccines are often used as supportive evidence towards other vaccines, such as COVID-19. ConclusionsOur findings underscore the importance of targeted public health interventions to restore and maintain vaccine confidence in Japan. While vaccine confidence has shown a slow increase, sustained efforts are necessary to secure long-term improvements. Confidence in one vaccine may positively influence perceptions of other vaccines. Addressing misinformation, reducing complacency, and enhancing vaccine accessibility are key strategies to improve uptake. Increased confidence in HPV vaccines appeared to have a positive influence on confidence in other vaccines, such as COVID-19. This study also demonstrates the utility of LLMs in offering a deeper understanding of public health attitudes. To effectively combat vaccine hesitancy and improve coverage, interventions must prioritize consistent communication, localized strategies, and an integrated approach to vaccine narratives.

6
Denoising Longitudinal Social Media for Pandemic Monitoring

Lin, S.; Garay, L.; Hua, Y.; Guo, Z.; Xu, X.; Yang, J.

2024-06-30 public and global health 10.1101/2024.06.29.24309690 medRxiv
Top 0.1%
45.2%
Show abstract

ObjectiveCurrent studies leveraging social media data for disease monitoring face challenges like noisy colloquial language and insufficient tracking of user disease progression in longitudinal data settings. This study aims to develop a pipeline for collecting, cleaning, and analyzing large-scale longitudinal social media data for disease monitoring, with a focus on COVID-19 pandemic. Materials and MethodsThis pipeline initiates by screening COVID-19 cases from tweets spanning February 1, 2020, to April 30, 2022. Longitudinal data is collected for each patient, two months before and three months after self-reporting. Symptoms are extracted using Name Entity Recognition (NER), followed by denoising with a combination of Graph Convolutional Network (GCN) and Bidirectional Encoder Representations from Transformers (BERT) model to retain only User Symptom Mentions (USM). Subsequently, symptoms are mapped to standardized medical concepts using the Unified Medical Language System (UMLS). Finally, this study conducts symptom pattern analysis and visualization to illustrate temporal changes in symptom prevalence and co-occurrence. ResultsThis study identified 191,096 self-reported COVID-19-positive cases from COVID-19-related tweets and retrospectively collected 811,398,280 historical tweets, of which 2,120,964 contained symptoms information. After denoising, 39% (832,287) of symptom-sharing tweets reflected user-related mentions. The trained USM model achieved an F1 score of 0.926. Further analysis revealed a higher prevalence of upper respiratory tract symptoms during the Omicron period compared to the Delta and wild-type periods. Additionally, there was a pronounced co-occurrence of lower respiratory tract and nervous system symptoms in the wild-type strain and Delta variant. ConclusionThis study established a robust framework for pandemic monitoring via social media, integrating denoising of user-related symptom mentions and longitudinal data. The findings underscore the importance of denoising procedures in revealing accurate prevalence trends, thereby minimizing biases in symptom analysis.

7
Analyzing Conspiratorial Content Across Singapore-Based Telegram Groups

Goyal, A.; Ligo, V. A. C.; Cheung, L. Y.; Lee, R. K.-W.; Saha, K.; Tandoc, E. C.; Kumar, N.

2025-07-17 public and global health 10.1101/2025.07.15.25331450 medRxiv
Top 0.1%
41.5%
Show abstract

Telegram has emerged as a key platform for the circulation of conspiratorial narratives. We examine conspiratorial discourse within Singapore-based Telegram groups from 2021-2025. We analyze over 10 million words from three Telegram groups. We developed a logistic regression classifier to detect conspiratorial content, achieving an F1 score of 0.74 and expert-validated labeling accuracy of 72%. Topic models indicated dominant themes centered around elite control, vaccine risks, and globalist agendas. While most users rarely posted conspiratorial content, a small, highly active minority accounted for most of such messages. These users frequently forwarded messages across multiple groups, amplifying the spread of content with short but intense lifecycles (mean lifespan=6.8 days). Network analysis showed that users typically joined multiple groups in rapid succession and that conspiratorial messages traveled across groups within weeks. We underscore the importance of user-centric monitoring, time-sensitive interventions, and platform-specific models for content detection.

8
Twitter activity about treatments during the COVID-19 pandemic: case studies of remdesivir, hydroxychloroquine, and convalescent plasma.

Hamamsy, T. C.; Bonneau, R.

2020-07-13 public and global health 10.1101/2020.06.18.20134668 medRxiv
Top 0.1%
40.1%
Show abstract

Since the COVID-19 pandemic started, the public has been eager for news about promising treatments, and social media has played a large role in information dissemination. In this paper, our objectives are to characterize the public discussion of treatments on Twitter, and demonstrate the utility of these discussions for public health surveillance. We pulled tweets related to three promising COVID-19 treatments (hydroxychloroquine, remdesivir and convalescent plasma), between the dates of February 28th and May 22nd using the Twitter public API. We characterize treatment tweet trends over this time period. Most major tweet/retweet/sentiment trends correlated to public announcement made by the white house and/or to new clinical trial evidence about treatments. Most of the websites people shared in treatment-related tweets were non-scientific media sources that leaned conservative. Hydroxychloroquine was the most discussed treatment on Twitter, and over 10% of hydroxychloroquine tweets mentioned an adverse drug reaction. There is a gap between the public attention/discussion around COVID-19 treatments and their evidence. Twitter data can and should be used for public health surveillance during this pandemic, as it is informative for monitoring adverse drug reactions, especially as many people avoid going to hospitals/doctors.

9
Unmasking the conversation on masks: Natural language processing for topical sentiment analysis of COVID-19 Twitter discourse

Sanders, A.; White, R.; Severson, L.; Ma, R.; McQueen, R.; Alcanatara Paulo, H. C.; Zhang, Y.; Erickson, J. S.; Bennett, K. P.

2020-09-01 health informatics 10.1101/2020.08.28.20183863 medRxiv
Top 0.1%
39.7%
Show abstract

In this exploratory study, we scrutinize a database of over one million tweets collected from March to July 2020 to illustrate public attitudes towards mask usage during the COVID-19 pandemic. We employ natural language processing, clustering and sentiment analysis techniques to organize tweets relating to mask-wearing into high-level themes, then relay narratives for each theme using automatic text summarization. In recent months, a body of literature has highlighted the robustness of trends in online activity as proxies for the sociological impact of COVID-19. We find that topic clustering based on mask-related Twitter data offers revealing insights into societal perceptions of COVID-19 and techniques for its prevention. We observe that the volume and polarity of mask-related tweets has greatly increased. Importantly, the analysis pipeline presented may be leveraged by the health community for qualitative assessment of public response to health intervention techniques in real time.

10
Impact of a Social Media Derived Digital Self Management Platform on Population Level Irritable Bowel Syndrome Emergency Utilization: A Controlled Interrupted Time Series Analysis Using South Korean National Health Insurance Data

Park, J.-H.; Lim, A.

2026-03-23 health informatics 10.64898/2026.03.20.26348871 medRxiv
Top 0.1%
39.7%
Show abstract

BackgroundIrritable bowel syndrome (IBS) contributes disproportionately to gastrointestinal-related emergency department (ED) utilization in South Korea, yet evidence on population-level interventions informed by patient-generated digital discourse remains limited. Recent social media analyses have identified dominant thematic concerns among IBS patients, including dietary triggers, symptom management, psychosocial burden, and information-seeking, suggesting actionable targets for digital self-management tools. ObjectiveTo evaluate the population-level impact of the Jang Geongang (, "Gut Health") digital self-management platform, whose content architecture was informed by topic modeling of IBS-related social media discourse, on IBS-attributed ED visits and unplanned hospitalizations, using a controlled interrupted time series (CITS) design. MethodsWe analyzed monthly aggregate claims data from South Koreas National Health Insurance Service (NHIS) spanning January 2018 to December 2024 (84 monthly observations). The Jang Geongang platform was launched in four pilot metropolitan areas (Seoul, Incheon, Daejeon, Gwangju) in July 2021, with eight non-pilot metropolitan areas serving as concurrent controls. Segmented regression with Newey-West heteroskedasticity and autocorrelation consistent (HAC) standard errors was used to estimate changes in level and trend of IBS-attributed ED visits per 100,000 insured population. Sensitivity analyses included autoregressive integrated moving average (ARIMA) transfer function models, varying pre-intervention windows, and leave-one-out control exclusion. ResultsThe CITS model estimated an immediate level change of -3.42 IBS-attributed ED visits per 100,000 (95% CI: -5.18 to -1.66, p < 0.001) following platform launch, and a change in monthly trend of -0.19 visits per 100,000 per month (95% CI: -0.31 to -0.07, p = 0.003), compared to control areas. By December 2024, the cumulative estimated reduction was 10.5 ED visits per 100,000 (23.8% relative reduction). Effects were concentrated in younger adults (19-39 years; level change: -5.14, p < 0.001) and IBS-D subtype visits (level change: -4.87, p < 0.001). ARIMA transfer function models corroborated these findings (immediate impact: -3.28, p = 0.001). Unplanned hospitalizations showed a smaller but significant reduction (level change: -0.84 per 100,000, p = 0.018). ConclusionsA digital self-management platform designed using social media derived IBS patient discourse insights was associated with sustained population-level reductions in IBS-attributed emergency utilization. Controlled interrupted time series analysis provides robust evidence for the public health impact of translating social media analytics into scalable digital health interventions.

11
COVID-19 vaccine perceptions: An observational study on Reddit

Kumar, N.; Corpus, I.; Hans, M.; Harle, N.; Yang, N.; McDonald, C.; Sakai, S. N.; Janmohamed, K. A.; Tang, W.; Schwartz, J. L.; Jones-Jang, S. M.; Saha, K.; Memon, S. A.; Bauch, C.; Chaudhury, M. D.; Papakyriakopoulos, O.; Tucker, J. D.; Goyal, A.; Tyagi, A.; Khoshnood, K.; Omer, S.

2021-04-13 public and global health 10.1101/2021.04.09.21255229 medRxiv
Top 0.1%
34.6%
Show abstract

ObjectivesAs COVID-19 vaccinations accelerate in many countries, narratives skeptical of vaccination have also spread through social media. Open online forums like Reddit provide an opportunity to quantitatively examine COVID-19 vaccine perceptions over time. We examine COVID-19 misinformation on Reddit following vaccine scientific announcements. MethodsWe collected all posts on Reddit from January 1 2020 - December 14 2020 (n=266,840) that contained both COVID-19 and vaccine-related keywords. We used topic modeling to understand changes in word prevalence within topics after the release of vaccine trial data. Social network analysis was also conducted to determine the relationship between Reddit communities (subreddits) that shared COVID-19 vaccine posts, and the movement of posts between subreddits. ResultsThere was an association between a Pfizer press release reporting 90% efficacy and increased discussion on vaccine misinformation. We observed an association between Johnson and Johnson temporarily halting its vaccine trials and reduced misinformation. We found that information skeptical of vaccination was first posted in a subreddit (r/Coronavirus) which favored accurate information and then reposted in subreddits associated with antivaccine beliefs and conspiracy theories (e.g. conspiracy, LockdownSkepticism). ConclusionsOur findings can inform the development of interventions where individuals determine the accuracy of vaccine information, and communications campaigns to improve COVID-19 vaccine perceptions. Such efforts can increase individual- and population-level awareness of accurate and scientifically sound information regarding vaccines and thereby improve attitudes about vaccines. Further research is needed to understand how social media can contribute to COVID-19 vaccination services. FundingStudy was funded by the Yale Institute for Global Health and the Whitney and Betty MacMillan Center for International and Area Studies at Yale University. The funding bodies had no role in the design, analysis or interpretation of the data in the study.

12
Social Media Reveals Psychosocial Effects of the COVID-19 Pandemic

Saha, K.; Torous, J.; Caine, E. D.; De Choudhury, M.

2020-10-26 psychiatry and clinical psychology 10.1101/2020.08.07.20170548 medRxiv
Top 0.1%
34.2%
Show abstract

BackgroundThe novel coronavirus disease 2019 (COVID-19) pandemic has caused several disruptions in personal and collective lives worldwide. The uncertainties surrounding the pandemic have also led to multi-faceted mental health concerns, which can be exacerbated with precautionary measures such as social distancing and self-quarantining, as well as societal impacts such as economic downturn and job loss. Despite noting this as a "mental health tsunami," the psychological effects of the COVID-19 crisis remains unexplored at scale. Consequently, public health stakeholders are currently limited in identifying ways to provide timely and tailored support during these circumstances. ObjectiveOur work aims to provide insights regarding peoples psychosocial concerns during the COVID-19 pandemic by leveraging social media data. We aim to study the temporal and linguistic changes in symptomatic mental health and support-seeking expressions in the pandemic context. MethodsWe obtain ~60M Twitter streaming posts originating from the U.S. from March, 24 - May, 25, 2020, and compare these with ~40M posts from a comparable period in 2019 to causally attribute the effect of COVID-19 on peoples social media self-disclosure. Using these datasets, we study peoples self-disclosure on social media in terms of symptomatic mental health concerns and expressions seeking support. We employ transfer learning classifiers that identify the social media language indicative of mental health outcomes (anxiety, depression, stress, and suicidal ideation) and support (emotional and informational support). We then examine the changes in psychosocial expressions over time and language, comparing the 2020 and 2019 datasets. ResultsWe find that all of the examined psychosocial expressions have significantly increased during the COVID-19 crisis - mental health symptomatic expressions have increased by ~14%, and support seeking expressions have increased by ~5%, both thematically related to COVID-19. We also observe a steady decline and eventual plateauing in these expressions during the COVID-19 pandemic, which may have been due to habituation or due to supportive policy measures enacted during this period. Our language analyses highlight that people express concerns that are contextually related to the COVID-19 crisis. ConclusionsWe studied the psychosocial effects of the COVID-19 crisis by using social media data from 2020, finding that peoples mental health symptomatic and support-seeking expressions significantly increased during the COVID-19 period as compared to similar data from 2019. However, this effect gradually lessened over time, suggesting that people adapted to the circumstances and their "new normal". Our linguistic analyses revealed that people expressed mental health concerns regarding personal and professional challenges, healthcare and precautionary measures, and pandemic-related awareness. This work shows the potential to provide insights to mental healthcare and stakeholders and policymakers in planning and implementing measures to mitigate mental health risks amidst the health crisis.

13
Exploring the Link Between Cancer Information Complexity and Understanding Medical Statistics in Online Health Information Seeking: Insights from Health Information National Trends Survey (HINTS)

CHAKRABORTY, A.; Das, S.; Phyo, M.

2026-03-20 health informatics 10.64898/2026.03.18.26348735 medRxiv
Top 0.1%
34.1%
Show abstract

Introduction: Understanding the factors influencing perceptions of cancer-related information is crucial for improving public health communication. This study explores the association between perceived difficulty in understanding information related to cancer (Cancer info Hard to Understand) and concerns about the quality of cancer-related information (Concern about Cancer Info Quality) with the extent of difficulty in comprehending medical statistics information (Understanding Medical Statistics). Methods: Data came from the 2022 Health Information National Trends Survey (HINTS). The cross-sectional study included 1972 participants with a response rate of 67.36% for Cancer info Hard to Understand, and 65.31% for Concern about Cancer Info Quality. We investigated the effect of Understanding Medical Statistics on Cancer info Hard to Understand, and Concern about Cancer Info Quality using univariate and multivariable logistic regression models with survey weights. The multivariable logistic regression model was adjusted for age, gender, ethnicity, marital status, education level, employment history, confidence in internet health resources, and social media. The chi-square test was used to measure the association between the predictors and the outcome. Results: Individuals finding medical statistics hard to understand were more likely to be concerned regarding the quality of the cancer-related information (AOR=1.74, 95% CI: [1.20, 2.52]) and also found cancer-related information difficult to comprehend (AOR=1.89, 95% CI: [1.19, 3.00]). Also, the influence of social media on health information seeking was significantly associated with Concern about Cancer Info Quality (AOR=2.24; 95% CI: [1.33, 3.76]), and Cancer info Hard to Understand (AOR=2.84; 95% CI: [1.61, 5.03]). Conclusion: This study highlights the critical role of understanding medical statistics in shaping perceptions of cancer-related information. From an epidemiological perspective, enhancing statistical literacy is essential for making informed health decisions, addressing health disparities, and designing effective, targeted cancer communication strategies.

14
A Chronological and Geographical Analysis of Personal Reports of COVID-19 on Twitter

Klein, A.; Magge, A.; O'Connor, K.; Cai, H.; Weissenbacher, D.; Gonzalez-Hernandez, G.

2020-04-22 health informatics 10.1101/2020.04.19.20069948 medRxiv
Top 0.1%
33.2%
Show abstract

The rapidly evolving outbreak of COVID-19 presents challenges for actively monitoring its spread. In this study, we assessed a social media mining approach for automatically analyzing the chronological and geographical distribution of users in the United States reporting personal information related to COVID-19 on Twitter. The results suggest that our natural language processing and machine learning framework could help provide an early indication of the spread of COVID-19.

15
The impact of technology systems and professional support in digital mental health interventions: a secondary meta-analysis

Sasseville, M.; LeBlanc, A.; Tchuente, J.; Boucher, M.; Dugas, M.; Mbemba, G.; Barony, R.; Chouinard, M.-C.; Beaulieu, M.; Beaudet, N.; Skidmore, B.; Cholette, P.; Aspiros, C.; Larouche, A.; Chabot, G.; Gagnon, M.-P.

2021-04-19 health informatics 10.1101/2021.04.12.21255333 medRxiv
Top 0.1%
32.7%
Show abstract

BackgroundA rapid review of systematic reviews was conducted to assess the effectiveness of digital mental health interventions for people with a chronic disease. Although it provided an overview of the evidence, it offered limited understanding of ethe types of interventions that were the most effective. The aim of this study was to perform a meta-analysis of primary studies identified in this rapid review of systematic reviews by focusing on the needs of knowledge users. MethodsThis secondary meta-analysis follows a rapid review of systematic reviews, a virtual workshop with knowledge users to identify research questions and a modified Delphi study to guide research methods. We conducted a secondary analysis of the primary studies identified in the rapid review. Two reviewers independently screened the titles and abstracts and applied inclusion criteria: RCT design using a digital mental health intervention in a population of adults with another chronic condition, published after 2010 in French or English, and including an outcome measurement of anxiety or depression. Results708 primary studies were extracted from the systematic reviews and 84 primary studies met the inclusion criteria Digital mental health interventions were significantly more effective than in-person care for both anxiety and depression outcomes. Online messaging was the most effective technology to improve anxiety and depression scores; however, all technology types were effective. Interventions partially supported by healthcare professionals were more effective than self-administered. ConclusionsWhile our meta-analysis identifies digital interventions characteristics that are more effective, all technologies and levels of support can be used considering implementation context and population. Review registrationThe protocol for this review is registered in the National Collaborating Centre for Methods and Tools (NCCMT) COVID-19 Rapid Evidence Service (ID 75).

16
Trend and co-occurrence network study of symptoms through social media: an example of COVID-19

Wu, J.; Wang, L.; Hua, Y.; Li, M.; Zhou, L.; Bates, D. W.; Yang, J.

2022-09-29 public and global health 10.1101/2022.09.28.22280462 medRxiv
Top 0.1%
32.0%
Show abstract

ImportanceCOVID-19 is a multi-organ disease with broad-spectrum manifestations. Clinical data-driven research can be difficult because many patients do not receive prompt diagnoses, treatment, and follow-up studies. Social medias accessibility, promptness, and rich information provide an opportunity for large-scale and long-term analyses, enabling a comprehensive symptom investigation to complement clinical studies. ObjectivePresent an efficient workflow to identify and study the characteristics and co-occurrences of COVID-19 symptoms using social media. Design, Setting, and ParticipantsThis retrospective cohort study analyzed 471,553,966 COVID-19-related tweets from February 1, 2020, to April 30, 2022. A comprehensive lexicon of symptoms was used to filter tweets through rule-based methods. 948,478 tweets with self-reported symptoms from 689,551 Twitter users were identified for analysis. Main Outcomes and MeasuresThe overall trends of COVID-19 symptoms reported on Twitter were analyzed (separately by the Delta strain and the Omicron strain) using weekly new numbers, overall frequency, and temporal distribution of reported symptoms. A co-occurrence network was developed to investigate relationships between symptoms and affected organ systems. ResultsThe weekly quantity of self-reported symptoms has a high consistency (0.8528, P<0.0001) and one-week leading trend (0. 8802, P<0.0001) with new infections in four countries. We grouped 201 common symptoms (mentioned [&ge;] 10 times) into 10 affected systems. The frequency of symptoms showed dynamic changes as the pandemic progressed, from typical respiratory symptoms in the early stage to more musculoskeletal and nervous symptoms at later stages. When comparing symptoms reported during the Delta strain versus the Omicron variant, significant changes were observed, with dropped odd ratios of coma (95%CI 0.55-0.49, P<0.01) and anosmia (95%CI, 0.6-0.56), and more pain in the throat (95%CI, 1.86-1.96) and concentration problems (95%CI, 1.58-1.70). The co-occurrence network characterizes relationships among symptoms and affected systems, both intra-systemic, such as cough and sneezing (respiratory), and inter-systemic, such as alopecia (integumentary) and impotence (reproductive). Conclusions and RelevanceWe found dynamic COVID-19 symptom evolution through self-reporting on social media and identified 201 symptoms from 10 affected systems. This demonstrates that social medias prevalence trends and co-occurrence networks can efficiently identify and study public health problems, such as common symptoms during pandemics. Key pointsO_ST_ABSQuestionsC_ST_ABSWhat are the epidemic characteristics and relationships of COVID-19 symptoms that have been extensively reported on social media? FindingsThis retrospective cohort study of 948,478 related tweets (February 2020 to April 2022) from 689,551 users identified 201 self-reported COVID-19 symptoms from 10 affected systems, mitigating the potential missing information in hospital-based epidemiologic studies due to many patients not being timely diagnosed and treated. Coma, anosmia, taste sense altered, and dyspnea were less common in participants infected during Omicron prevalence than in Delta. Symptoms that affect the same system have high co-occurrence. Frequent co-occurrences occurred between symptoms and systems corresponding to specific disease progressions, such as palpitations and dyspnea, alopecia and impotence. MeaningTrend and network analysis in social media can mine dynamic epidemic characteristics and relationships between symptoms in emergent pandemics.

17
Quality of Chronic Disease Related Health Videos Across Social Media Platforms: A Systematic Review and Meta analysis

Liu, R.; Xu, Y.; Zhang, M.; Li, Y.; Wang, X.; Huang, C.

2026-07-15 public and global health 10.64898/2026.07.14.26358040 medRxiv
Top 0.1%
31.9%
Show abstract

Background: Social media videos have become one of the major sources of health information for individuals living with chronic diseases. Although numerous cross-sectional studies have evaluated the quality of health-related videos across different platforms, the overall quality of chronic disease-related videos and the determinants underlying quality variation remain unclear. Objective: To systematically evaluate the quality of chronic disease-related health videos across major global and Chinese social media platforms and to identify potential determinants of video quality using multivariable meta-regression. Methods: This systematic review and meta-analysis searched PubMed, Embase, and Web of Science from database inception to April 30, 2026, for cross-sectional studies evaluating Chinese-and English-language health videos. Scores from the DISCERN instrument, the Global Quality Scale (GQS), and the Journal of the American Medical Association (JAMA) benchmark criteria were standardized to a 0-100 scale and quantitatively synthesized using random-effects models. Prespecified subgroup analyses and multivariable meta-regression were conducted to explore potential sources of heterogeneity, including platform region, platform type, disease category, video duration, professional background of content creators, and audience engagement. Results: A total of 88 studies involving 18,688 videos were included. Overall methodological quality was suboptimal, with a pooled standardized DISCERN score of 50.61 (95% CI, 48.04-53.18), accompanied by substantial between-study heterogeneity (I2 = 99.3%). Videos hosted on international platforms achieved significantly higher quality scores than those on Chinese platforms (54.68 vs. 48.21; P = 0.008). Multivariable meta-regression demonstrated that conventional predictors-including video duration, the proportion of physician creators, and audience engagement-were not independently associated with video quality (P > 0.05). Importantly, the final model explained only 9.06% of the between-study heterogeneity (R2 = 9.06%), indicating that conventional content- and creator-level characteristics account for only a small proportion of the observed variability in video quality. Conclusions: Traditional predictors, including creator professionalism, video duration, disease category, and audience engagement, have limited ability to explain variation in the quality of online health videos. Although platform region emerged as the only significant moderator, the multivariable model explained only a small fraction of the observed heterogeneity, suggesting that the principal determinants of health information quality remain largely unexplained.

18
Divide in Vaccine Belief in COVID-19 Conversations: Implications for Immunization Plans

Tyagi, A.; Carley, K. M.

2020-07-29 health policy 10.1101/2020.07.23.20160887 medRxiv
Top 0.1%
31.6%
Show abstract

The development of a viable COVID-19 vaccine is a work in progress, but the success of the immunization campaign will depend upon public acceptance. In this paper, we classify Twitter users in COVID-19 discussion into vaccine refusers (anti-vaxxers) and vaccine adherers (vaxxers) communities. We study the divide between anti-vaxxers and vaxxers in the context of whom they follow. More specifically, we look at followership of 1) the U.S. Congress members, 2) four major religions (Christianity, Hinduism, Judaism and Islam), 3) accounts related to the healthcare community, and 4) news media accounts. Our results indicate that there is a partisan divide between vaxxers and anti-vaxxers. We find a religious community with a higher than expected fraction of anti-vaxxers. Further, we find that the variance of vaccine belief within the news media accounts operated by Russian and Iranian governments is higher compared to news media accounts operated by other governments. Finally, we provide messaging and policy implications to inform the COVID-19 vaccine and future vaccination plans.

19
Early detection of fraudulent COVID-19 products from Twitter chatter

Sarker, A.; Lakamana, S.; Liao, R.; Abbas, A.; Yang, Y.-C.; Al-Garadi, M.

2022-05-11 public and global health 10.1101/2022.05.09.22274776 medRxiv
Top 0.1%
31.2%
Show abstract

Social media have served as lucrative platforms for misinformation and for promoting fraudulent products for the treatment, testing and prevention of COVID-19. This has resulted in the issuance of many warning letters by the United States Food and Drug Administration (FDA). While social media continue to serve as the primary platform for the promotion of such fraudulent products, they also present the opportunity to identify these products early by employing effective social media mining methods. In this study, we employ natural language processing and time series anomaly detection methods for automatically detecting fraudulent COVID-19 products early from Twitter. Our approach is based on the intuition that increases in the popularity of fraudulent products lead to corresponding anomalous increases in the volume of chatter regarding them. We utilized an anomaly detection method on streaming COVID-19-related Twitter data to detect potentially anomalous increases in mentions of fraudulent products. Our unsupervised approach detected 34/44 (77.3%) signals about fraudulent products earlier than the FDA letter issuance dates, and an additional 6/44 (13.6%) within a week following the corresponding FDA letters. Our proposed method is simple, effective and easy to deploy, and do not require high performance computing machinery unlike deep neural network-based methods.

20
Large language models for self-administered conversational vignette assessment of provider competencies: A pilot and validation study in Vietnam with automated LLM-powered transcript classification

Daniels, B.; Zhang, W.; Nguyen, H.; Duong, D.

2026-03-04 health economics 10.64898/2026.03.02.26347479 medRxiv
Top 0.1%
31.1%
Show abstract

We developed and validated a self-administered clinical vignette platform powered by a large language model (LLM), deployed through a SurveyCTO web survey, to measure primary health care provider competencies in Vietnam. In a pilot focus group, nine physicians rated LLM-simulated patient interactions as realistic (mean 3.78/5) and user-friendly. In the validation phase, 22 providers completed 132 vignette interactions across ten clinical scenarios in Vietnamese. Essential diagnostic checklist scores (human-coded from translated transcripts) correlated with expert clinician evaluations (Pearsons{rho} = 0.55-0.60). LLM-automated coding of checklist items from translated English transcripts correlated reasonably with human coding ({rho} = 0.53), and coding directly from Vietnamese transcripts performed comparably ({rho} = 0.51), suggesting that a separate translation step may not be necessary. The total cost of 132 chatbot interactions was under USD 2. LLM-driven conversational vignettes represent a low-cost and scalable method for assessing provider competencies in respondents local language, eliminating the need for extensive enumeration staffs while preserving the open-ended format critical to vignette validity, and additionally introducing flexible feature extraction from transcripts using grading rubrics. The platform is open-source and designed for replication in other health system contexts. Author summaryMeasuring the clinical skills of healthcare providers is essential for improving the quality of care, but current survey methods are expensive and require trained enumerators to travel to health facilities in person. We developed a new approach that uses large language models (LLMs) - the technology behind tools like ChatGPT and Claude - to simulate patients in realistic clinical conversations that healthcare providers can complete on their phones or laptops over the Internet in their own language. In Vietnam, we tested this tool with 31 physicians across ten clinical scenarios. Providers found the simulated patient conversations realistic and easy to use. We also tested whether LLMs could automatically score the conversations, which showed reasonable agreement with human scoring, and performed nearly as well when scoring directly from Vietnamese, without requiring a separate translation step. When we compared these results from our tool against holistic expert physician ratings of the same conversations, the scores agreed well, suggesting that automatic transcript grading based on rubrics produces meaningful measures of clinical skill. This tool costs less than two US dollars for over a hundred consultations and required no in-person surveyors, making it potentially transformative for routine, large-scale monitoring of healthcare quality in resource-limited settings. The platform and code are openly available for adaptation.