Can Large Language Models Diagnose Primary Immunodeficiency from Patient-Described Symptoms?
Reteig, L. C.; Woloshin, S.; Maglione, P. J.; Farmer, J. R.; Ong, M.-S.
Show abstract
Patients with primary immunodeficiency (PID) often face prolonged diagnostic delays and may increasingly turn to large language models (LLMs) to interpret their symptoms during this period. We evaluated whether an LLM could recognize PID from symptom descriptions derived from interviews with 21 PID patients. In a prior study, we showed that GPT-4o identified PID in 96% of cases when prompted with physician-written patient histories (Rider et al., JACI, 2024). Here, when prompted with symptom descriptions in patients' own words, GPT-5 identified PID in only 7 cases (33%), although it more broadly suggested immune system issues in 18 cases (81%). The gap between these findings indicates that LLMs are sensitive to the language and framing of symptom descriptions, performing substantially worse when patients describe their own symptoms in everyday language than when clinicians summarize patient histories in structured medical terms. This study underscores the need to carefully evaluate how LLMs are used in patient-facing applications.
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Loss of function mutation in ELF4 causes autoinflammatory and immunodeficiency disease in human 88%
- Early and rapid identification of COVID-19 patients with neutralizing type I-interferon auto-antibodies by an easily implementable algorithm 86%
- Long-term clinical, virological and immunological outcomes in patients hospitalized for COVID-19: antibody response predicts long COVID 86%
Similar papers in this journal
- Integrating circulating T follicular memory cells and autoantibody repertoires for characterization of autoimmune disorders 88%
- Cost-effectiveness of implementing objective diagnostic verification of asthma in the United States 87%
- Personal network inference identifies children at risk of recurrent wheezing and asthma 86%
Similar papers in this journal
- Clinical symptoms among ambulatory patients tested for SARS-CoV-2 87%
- High proportion of post-acute sequelae of SARS-CoV-2 infection in individuals 1-6 months after illness and association with disease severity in an outpatient telemedicine population 87%
- Sex and gender differences in COVID testing, hospital admission, presentation, and drivers of severe outcomes in the DC/Maryland region 87%
Similar papers in this journal
- Clinical characteristics and outcomes of COVID-19 breakthrough infections among vaccinated patients with systemic autoimmune rheumatic diseases 89%
- Temporal trends in COVID-19 outcomes among patients with systemic autoimmune rheumatic diseases: From the first wave to Omicron 89%
- IL-1 and IL-6 inhibitor hypersensitivity link to common HLA-DRB1*15 alleles 89%
Similar papers in this journal
- Differential antibody production by symptomatology in SARS-CoV-2 convalescent individuals 89%
- Clinical and laboratory features of COVID-19 illness and outcomes in immunocompromised individuals during the first pandemic wave in Sydney, Australia 89%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.