Back

Initial development and pragmatic clinical validation of a static disease severity instrument for pyoderma gangrenosum: Investigator Global Assessment for PG (IGAPg).

Jacobson, M. J.; Ng, J.; Morales Leon, L.; Keller, J.; Marzano, A.; Huang, W.; Kelly, R.; Mostaghimi, A.; Ortega-Loayza, A.

2026-01-01 dermatology
10.64898/2025.12.26.25342857 medRxiv
Show abstract

BackgroundPyoderma gangrenosum (PG) is a rare neutrophilic ulcerative dermatosis with no FDA-approved therapies and limited validated outcome measures. Investigators Global Assessments (IGAs) are widely used in dermatology but often lack objective criteria, robust validation, and comparability across studies. There remains a critical need for a PG-specific, standardized, and validated severity instrument. MethodsWe developed and conducted initial validation of the Investigator Global Assessment for Pyoderma Gangrenosum (IGAPg(C)), a novel PG-specific IGA. An international multistakeholder panel guided development, with a core team designing the instrument based on prior research identifying key objective clinical features of PG severity. The IGAPg(C) incorporates ulcer depth, drainage, discoloration, and undermining, with ulcer location and extent used to resolve indeterminate cases. Standardized rater training materials were created. Construct validity was assessed in 36 patients evaluated by an expert PG clinician, with correlations to Patient Global Assessment (PGA), Skindex-Mini, and numeric rating scales (NRS) for 24-hour and 7-day pain. Inter-rater reliability was evaluated in a subset of 26 patients assessed independently by five raters using a two-way random-effects intraclass correlation coefficient (ICC [2,1]) within a linear mixed-effects model. ResultIGAPg(C) scores demonstrated strong correlation with PGA when assessed by an expert PG dermatologist (Pearsons r = 0.73) and when averaged across all raters (r = 0.69). Moderate correlations were observed with Skindex-Mini (r = 0.49), 24-hour pain NRS (r = 0.48), and 7-day pain NRS (r = 0.52), consistent with differing constructs measured. Inter-rater reliability was high (ICC = 0.76). Raters reported the instrument to be comprehensive, comprehensible, and efficient to administer. Reliability and construct validity metrics are summarized in Table 1. O_TBL View this table: org.highwire.dtl.DTLVardef@ff22edorg.highwire.dtl.DTLVardef@4e33aaorg.highwire.dtl.DTLVardef@21f9c0org.highwire.dtl.DTLVardef@93d13forg.highwire.dtl.DTLVardef@8e96fc_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 1.C_FLOATNO O_TABLECAPTIONConstruct validity and inter-rater reliability of the IGAPg morphological tool. Correlations between IGAPg scores and patient-reported outcomes (PGA, Skindex-Mini, and average pain) demonstrate moderate-to-strong construct validity. Values are shown for all raters, Rater 1 alone, and Raters 2-5. Inter-item PRO correlations confirm internal consistency. Pearson correlations between Rater 1 and Raters 2-5 are reported with shared patient counts and p-values (p < 0.05 was considered significant). ICC(2,1) indicated good inter-rater reliability with most score variability attributable to differences between patients with minimal variability due to rater differences and a moderate residual component. Inter-item patient reported outcomes correlations confirm difference in measured constructs. C_TABLECAPTION C_TBL ConclusionsDespite limitations including modest sample size and rater homogeneity, the IGAPg(C) demonstrated strong construct validity and high inter-rater reliability. Designed for dermatologists and trainees familiar with PG morphology, the IGAPg(C) represents a promising PG-specific outcome measure for clinical research and therapeutic trials. Future work will focus on refining training materials and expanding validation with structured patient-investigator engagement.

Published in Journal of the American Academy of Dermatology · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.