Back

Evaluating the Ruptured Arteriovenous Malformation Grading Scale (RAGS): A Reliability Study

George, J.; Rahm, S. P.; Boudreau, H.; Hale, A. T.; Atchley, T. J.; Laskay, N. M.; Schmalz, P. G.; Jones, J.; Liptrap, E. J.; Harrigan, M. R.; Fisher, W. S.; Estvez-Ordonez, D.

2025-06-28 neurology
10.1101/2025.06.27.25330410 medRxiv
Show abstract

BackgroundUnruptured arteriovenous malformations (AVMs) carry a 1% risk of annual risk of hemorrhage; however, risk of re-hemorrhage is significantly higher after an initial rupture. Ruptured AVMs can cause significant morbidity and mortality. The Ruptured Arteriovenous Malformation Grading Scale (RAGS) was developed to better predict outcomes in patients with ruptured AVMs. However, the reliability of this scale has yet to be confirmed to support its use in clinical practice. ObjectiveTo determine the intra- and inter-rater reliability of RAGS among those most likely to use it in clinical practice. MethodsA cross-sectional sample of 42 patients with ruptured AVMs was selected via retrospective review, and clinical vignettes were created. Five raters were chosen to assign a RAGS score to all 42 patients to assess the inter-rater reliability of RAGS. After two months, ten patients from the study sample were randomly selected to be re-rated to determine the intra-rater reliability of RAGS. ResultsThe overall agreement rate was 97.2% among all raters. The inter-rater reliability was found to be substantial when measured using Cohen/Congers Kappa (0.73, 95% confidence interval (CI) [0.63, 0.82]), Scott/Fleiss Kappa (0.72, 95% CI [0.62, 0.82]), Krippendorfs Alpha (0.73, 95% CI [.63, 95% CI [0.63, 0.82]), intraclass correlation coefficient (ICC) (0.78, 95% CI [0.68, 0.86]), and Kendalls W (0.79, 95% CI [0.68, 0.86]) and almost perfect using Gwets AC (0.90, 95% CI [0.88, 0.93]). The test-retest percent agreement was between 94.7% and 98.1% among raters. ConclusionsThe RAGS classification system is highly reliable and has substantial to near-perfect agreement among raters with different expertise levels and specialties. This study supports the potential use of RAGS in clinical practice across different institutions. The overall agreement rate was 97.2% among all raters. The inter-rater reliability was found to be substantial when measured using Cohen/Congers Kappa (0.73, 95% confidence interval (CI) [0.63, 0.82]), Scott/Fleiss Kappa (0.72, 95% CI [0.62, 0.82]), Krippendorfs Alpha (0.73, 95% CI [.63, 95% CI [0.63, 0.82]), intraclass correlation coefficient (ICC) (0.78, 95% CI [0.68, 0.86]), and Kendalls W (0.79, 95% CI [0.68, 0.86]) and almost perfect using Gwets AC (0.90, 95% CI [0.88, 0.93]). The test-retest percent agreement was between 94.7% and 98.1% among raters.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.