Back

Patient-Reported Challenges in Lymphoma Diagnosis: Analysis of Online Forum Narratives Using Artificial Intelligence

He, F.; Valderas, J. M.

2025-11-19 health systems and quality improvement
10.1101/2025.09.07.25335273 medRxiv
Show abstract

BackgroundLymphoma diagnosis remains challenging due to diverse subtypes and nonspecific presentations. While prior research focused primarily on clinical accuracy, how patients experience and describe these challenges remain understudied. This study systematically analyzed online patient narratives to investigate their perspectives on diagnostic difficulties. MethodsWe developed an artificial intelligence (AI) pipeline integrating DeepSeek large language model, optical character recognition, and transformer-embedded keyword clustering to analyse online narratives reporting diagnostic discrepancies or misdiagnosis from Chinas largest lymphoma forum for patients and caregivers (house086.com). The pipeline extracted patient demographics, timelines, diagnostic barriers, facilitators, outcomes, and AI-graded severity. External validation (n=400) against manually-derived labels assessed pipeline reliability. Multivariable logistic regression examined associations between barriers, facilitators, and the binarized severity scores. ResultsOver the study period (2011-2025), patients reporting diagnostic difficulties doubled, while the rate of severe outcomes declined. From 2016 narratives (median patient age 47; 59% family-authored), AI-assisted keyword taxonomy identified 7 diagnostic facilitators, 11 barriers, and 5 consequences. Psychological distress was the most common consequence of diagnostic challenges (90%). Clinician-related issues (91%) and case complexity (77%) were the most prevalent barriers, but inappropriate initial treatment conferred the greatest risk (OR 19.06, 11.30-32.17). Among facilitators, specialist input reduced severe outcomes by 40% (OR 0.60, 0.44-0.81), while peer networks (OR 0.62) and clinician expertise (OR 0.65) provided additional protection. ConclusionThis large-scale analysis of patient narratives identified factors underlying patient-perceived diagnostic difficulties in lymphoma. AI enables scalable analysis of patient-generated data, offering insights into targeted quality improvement and digital health interventions in patient safety.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.