Noise and neglect: Social-media signals expose attention gaps for dengue, chikungunya, lymphatic filariasis and kala-azar in Indias vector-borne NTDs
Biswal, D. K.; Konhar, R.; Lalsanga, J. K.
Show abstract
BackgroundNeglected tropical diseases (NTDs), including dengue, chikungunya, lymphatic filariasis, and kala-azar, pose significant public health burdens in India. Despite WHO recommendations for enhanced disease surveillance and targeted communication strategies, little is known about public perceptions and discussions of these diseases across digital platforms. Understanding these perceptions can guide evidence-based policy making and public health messaging. MethodsWe conducted an in silico analysis of publicly accessible social and news media data related to dengue, chikungunya, filariasis, and kala-azar in India from January 2019 to December 2023. YouTube comments and Google News headlines were systematically retrieved, pre-processed, and analyzed through sentiment analysis (VADER lexicon) and Latent Dirichlet Allocation (LDA) topic modeling. Facebook and Twitter data were not included due to API restrictions and their current subscription-based models, limiting free access even for research purposes. We visualized disease-specific digital attention in comparison to epidemiological burden and created chord, Sankey, and network diagrams to elucidate thematic and sentiment-based interactions. ResultsDengue dominated online attention, accounting for over 50% of total mentions, despite a comparable or lower disease burden than filariasis and chikungunya. Kala-azar received minimal online engagement, highlighting a critical awareness gap. Sentiment analysis revealed predominantly neutral-to-positive discourse, especially focused on treatments, preventive measures, and vaccination initiatives. Topic modeling highlighted recurrent themes, including public health campaigns, outbreak alerts, and community-based interventions. ConclusionsOur study presents a novel approach combining digital surveillance, sentiment analysis, and topic modeling to provide insights into public perceptions of NTDs in India. The observed mismatch between epidemiological burden and online attention underscores the need for strategic public health messaging, aligning with WHO recommendations for community engagement and tailored disease-awareness campaigns. This research provides a valuable tool for policymakers to enhance the effectiveness of communication strategies and improve targeted intervention planning for neglected tropical diseases in India. Author SummaryNeglected tropical diseases (NTDs)--including dengue, chikungunya, lymphatic filariasis and kala-azar--still afflict millions across India, yet the public conversation remains uneven. We examined more than 45 000 YouTube comments and 270 Google News reports posted between January 2019 and December 2023 to see how these four NTDs are discussed online. After automated text cleaning, VADER sentiment scoring and Latent Dirichlet Allocation topic modelling, we overlaid the resulting tone-and-topic maps on official disease-burden data. Dengue dominated the chatter, accounting for well over half of all references, whereas kala-azar, though still endemic, drew scarcely any notice. Overall sentiment skewed neutral-to-positive and focused largely on prevention, treatment and vaccine news. Interactive bubble maps, Sankey flows and chord diagrams vividly exposed the gulf between epidemiological need and digital attention. We could not analyse Facebook or Twitter because their new, pay-walled APIs make large-scale data collection prohibitively expensive for researchers, underscoring a growing obstacle for digital epidemiology. Our reproducible, low-cost workflow highlights which NTDs are being overlooked online, providing Indian health authorities with actionable evidence and supporting the World Health Organizations call for stronger community engagement in the fight against NTDs.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 92%
- Use of large language models as a scalable approach to understanding public health discourse 91%
- The MARC SE-Africa Dashboard: Joining Forces to Counteract Emerging Antimalarial Resistance in South and East Africa 91%
Similar papers in this journal
- An in-depth statistical analysis of the COVID-19 pandemic’s initial spread in the WHO African region 92%
- Maternal and perinatal health research during emerging and ongoing epidemic threats: a landscape analysis and expert consultation 91%
- SARS-CoV-2 infection in Africa: A systematic review and meta-analysis of standardised seroprevalence studies, from January 2020 to December 2021 91%
Similar papers in this journal
- Quantifying the online news media coverage of the COVID-19 pandemic 90%
- Predicting public take-up of digital contact tracing during the COVID-19 crisis: Results of a national survey 90%
- “Is this Herpes or Syphilis?”: Latent Dirichlet Allocation Analysis of Sexually Transmitted Disease-Related Reddit Posts During the COVID-19 Pandemic 90%
Similar papers in this journal
- Using Google Health Trends to investigate COVID-19 incidence in Africa 95%
- Integrating digital and field surveillance to complement efforts to manage epidemic diseases of livestock: African swine fever as a case study 91%
- Recommended distances for physical distancing during COVID-19 pandemics reveal cultural connections between countries 91%
Similar papers in this journal
- How does policy modelling work in practice? A global analysis on the use of epidemiological modelling in health crises 92%
- Gaps and Opportunities for Data Systems and Economics to Support Priority Setting for Climate-Sensitive Infectious Diseases in Sub-Saharan Africa: A Rapid Scoping Review 91%
- The Global Impact of COVID-19 on Tuberculosis: A Thematic Scoping Review, 2020-2023 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.