The Patients' Voice in Clostridioides difficile Infection: Large Language Model-Assisted Thematic Analysis of Patient Testimonials
Villafuerte-Galvez, J. A.; Noriega, M. A.; Cakir Colak, S.; Crawford, C. V.
Show abstract
Background. Clostridioides difficile infection (CDI) imposes a burden that extends well beyond the gastrointestinal tract, yet existing outcome measures only partially capture the patient experience. We used frontier large language models (LLMs) on patient and caregiver narratives at scale to describe how burden shifts with disease course. Methods. We analyzed 189 testimonials from the Peggy Lillis Foundation corpus, sorted into four cohorts with recurrence (r) and fulminant (f) severity as axes (rfCDI, fCDI, rCDI, non-rfCDI). Two independent LLMs coded eight thematic domains, four fulminant flags, thirteen emerging semantic fields, the dominant dimension, and narrative arcs. Two clinicians independently coded a subset for inter-rater reliability (PABAK, Gwet's AC1). Results. Treatment trajectory was the dominant theme in recurrent disease, whereas death and near-death dominated non-recurrent fulminant narratives. Psychological burden was near-universal in fulminant disease (98.0% in rfCDI, 97.2% in fCDI). Caregiver and bereavement content concentrated in fCDI (66.7%). Diagnostic failure was frequent across recurrent cohorts (47.6 - 56.1%). Bacteriotherapy tracked recurrence (60.2% rfCDI versus 5.6% fCDI). Financial, mental-health, and caregiver burdens were prominent and are currently unaddressed by guidelines. Human-human reliability was substantial (PABAK 0.79 for semantic fields, 0.76 for domains); arc coding was least reliable. Conclusions. Patient narratives reveal a course-dependent, multidimensional burden in CDI. Concrete gaps exist between what patients prioritize, what guidelines recommend, and what therapy access provides. Frontier-LLM coding, validated against clinicians, offers a reproducible route to translate these priorities into research, care, and policy.
Matching journals
The top 13 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical outcomes among patients infected with Omicron (B.1.1.529) SARS-CoV-2 variant in southern California 90%
- Attributes and predictors of Long-COVID: analysis of COVID cases and their symptoms collected by the Covid Symptoms Study App 89%
- Psychiatric disorders and self-harm across 26 adult cancers: cumulative burden, temporal variation, excess years of life lost and unnatural causes of deaths 88%
Similar papers in this journal
- Study protocol for the Innovative Support for Patients with SARS-COV-2 Infections Registry (INSPIRE): a longitudinal study of the medium and long-term sequelae of SARS-CoV-2 infection 90%
- Characteristics of Long Covid: findings from a social media survey 90%
- Healthcare presentations with self-harm and the association with COVID-19: an e-cohort whole-population-based study using individual-level linked routine electronic health records in Wales, UK, 2016 - March 2021 90%
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 93%
- Identifying clusters of people with Multiple Long-Term Conditions using Large Language Models: a population-based study 91%
- New Model, Old Risks? Sociodemographic Bias and Adversarial Hallucinations Vulnerability in GPT-5 91%
Similar papers in this journal
- Epidemiology and costs of post-sepsis morbidity, nursing care dependency, and mortality in Germany 90%
- The Impact of the “Muslim Ban” Executive Order on Healthcare Utilization in Minneapolis-St. Paul, Minnesota 89%
- Estimating the Burden of Influenza on Daily Activity at Population Scale Using Commercial Wearable Sensors 88%
Similar papers in this journal
- Synchronous Caregiving from Birth to Adulthood Tunes Humans' Social Brain 88%
- Characterizing Population-level Changes in Human Behavior during the COVID-19 Pandemic in the United States 87%
- Large-scale genomic study reveals robust activation of the immune system following advanced Inner Engineering meditation retreat. 86%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.