Back

Assessing Pain Catastrophizing Through Free-Text Responses: A Validation of Large Language Models

Lee, A.; You, D. S.; Dildine, T. C.

2026-07-24 pain medicine
10.64898/2026.07.22.26358710 medRxiv
Show abstract

Validated measures of pain catastrophizing primarily assess catastrophizing as a stable trait. However, emerging evidence suggests catastrophizing fluctuates with context, highlighting a need for ecologically valid methods to capture it. This study evaluated large language models (LLMs) as implicit markers of catastrophizing from free text responses from ninety-one adults with chronic pain receiving long-term opioid therapy (57.3 percent Female; mean age = 60.5 years). Patients completed baseline measures, including the trait pain catastrophizing scale (PCS), followed by a 10 minute writing task after random assignment to a negative, positive, or neutral pain-coping condition. State affect and pain were assessed before and after writing tasks and again after a cold pressor task (4 degrees C; <= 2 minutes). A state PCS followed the cold pressor task. Free text responses were analyzed using four LLMs (Claude Opus 4; GPT Mini 4o; Llama 4 Maverick; and Gemini 2.5 Pro). ANOVA based results supported discriminant validity, as all four LLM-derived pain catastrophizing scores differentiated negative from positive and neutral pain-coping conditions. Convergent validity was model dependent; only Gemini derived scores correlated with state catastrophizing (r = .22) and pain unpleasantness (r = .23). Divergent validity was mixed. LLM derived scores were unrelated to pain intensity, but Gemini and Claude derived scores showed small correlations with trait PCS (rs = .21; 28, respectively). All LLM-derived scores also correlated with negative affect (rs range = .29 - .41), comparable in magnitude to state PCS, suggesting limited specificity. These findings provide preliminary evidence that certain LLMs may serve as implicit markers of state pain catastrophizing, but further study is needed.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.