Back

Multi-Specialty Expert Physician Identification of Extranodal Extension in Computed Tomography Scans of Oropharyngeal Cancer Patients: Prospective Blinded Human Inter-Observer Performance Evaluation

Sahin, O.; Wahid, K. A.; Taku, N.; He, R.; Naser, M. A.; Mohamed, A. S. R.; Makitie, A.; Kann, B. H.; Kaski, K.; Sahlsten, J.; Jaskari, J.; Amit, M.; Chen, M. M.; Chronowski, G. M.; Diaz, E. M.; Garden, A. S.; Goepfert, R. P.; Guenette, J. P.; Gunn, G. B.; Hirvonen, J.; Hoebers, F.; Guha-Thakurta, N.; Johnson, J.; Kaya, D.; Khanpara, S. D.; Nyman, K.; Lai, S. Y.; Lango, M.; Learned, K. O.; Lee, A.; Lewis, C. M.; Maniakas, A.; Moreno, A. C.; Myers, J. N.; Phan, J.; Pytynia, K. B.; Rosenthal, D. I.; Sandulache, V.; Schellingerhout, D.; Shah, S. J.; Sikora, A. G.; Wintermark, M.; Fuller, C. D.

2023-02-26 oncology
10.1101/2023.02.25.23286432 medRxiv
Show abstract

ImportanceExtranodal extension (pENE) is a critical prognostic factor in oropharyngeal cancer (OPC) that drives therapeutic disposition. Determination of pENE from radiological imaging has been associated with high inter-observer variability. However, the impact of clinician specialty on human observer performance of imaging-detected extranodal extension (iENE) remains poorly understood. ObjectiveTo characterize the impact of clinician specialty on the accuracy of pre-operative iENE in human papillomavirus-positive (HPV+) OPC using computed tomography (CT) images. Design, Setting, and ParticipantsThis prospective observational human performance study analyzed pre-therapy CT images from 24 HPV+ OPC patients, with duplication of 6 scans (n=30) of which 21 were pathologically confirmed pENE. Thirty-four expert observers, including 11 radiologists, 12 surgeons, and 11 radiation oncologists, independently assessed these scans for iENE and reported human-detected radiologic criteria and observer confidence. Main Outcomes and MeasuresThe primary outcomes included accuracy, sensitivity, specificity, area under the receiver operating characteristic curve (AUC), and Brier score for each physician, compared to ground-truth pENE. The significance of radiographic signs for prediction of pENE were determined through logistic regression analysis. Fleiss kappa measured interobserver agreement, and Hanley-MacNeil AUC discrimination testing. ResultsMedian accuracy across all specialties was 0.57 (95%CI 0.39 to 0.73), with no specialty showing discriminate performance greater than random estimation (median AUC 0.64, 95%CI 0.44 to 0.83). Significant differences between radiologists and surgeons in Brier scores (0.33 vs. 0.26, p < 0.01), radiation oncologists and surgeons in sensitivity (0.48 vs. 0.69, p > 0.1), and radiation oncologists and radiologists/surgeons in specificity (0.89 vs. 0.56, p > 0.1). Indistinct capsular contour and nodal necrosis were significant predictors of correct pENE status among all specialties. Interobserver agreement was weak for all the radiographic criteria, regardless of specialty ({kappa}<0.6). Conclusions and RelevanceMultiobserver testing shows physician discrimination of HPV+OPC pENE on pre-operative CT remains non-different than blind guessing, with high inter-rater variability and low diagnostic accuracy, regardless of clinician specialty. While minor differences in diagnostic performance among specialties are noted, they do not significantly affect the overall poor agreement and discrimination rates observed. The findings underscore the need for further research into automated detection systems or enhanced imaging techniques to improve the accuracy and reliability of iENE assessments in clinical practice. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=112 SRC="FIGDIR/small/23286432v2_ufig1.gif" ALT="Figure 1"> View larger version (38K): org.highwire.dtl.DTLVardef@2eefa7org.highwire.dtl.DTLVardef@177f053org.highwire.dtl.DTLVardef@142fcc6org.highwire.dtl.DTLVardef@e14eb0_HPS_FORMAT_FIGEXP M_FIG Visual Abstract C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.