Back

Machine Learning Based Classification of Aggressive and Malignant Renal Tumors from Multimodal Data

Aminy, M.; Gala, T.; Dasgupta, A.; Cen, S. Y.; Jogi, P. S.; Gill, I.; Duddalwar, V.; Oberai, A.

2025-02-06 radiology and imaging
10.1101/2025.02.04.25321687 medRxiv
Show abstract

1PurposeThis study aimed to develop and evaluate a machine learning pipeline using multiphase contrast-enhanced CT images and clinical data to classify renal tumors as benign, malignant-indolent, or malignant-aggressive, while assessing the contribution of each data source to the classification. MethodsIn this retrospective study, 448 patients (mean age: 60.7{+/-}12.6 years, 306 male, 142 female) who underwent nephrectomy and preoperative CECT between June 2008 and July 2018 were included. Tumors were histologically categorized as benign-indolent, malignant-indolent, or malignant-aggressive. Self-supervised feature extraction converted 4-phase CECT images into 512 real-valued features, combined with clinical data and tumor size for classification. Two machine learning classifiers, random forest (RF) and multi-layer perceptron (MLP), were used to predict tumor type. Nested five-fold cross-validation was employed for hyperparameter tuning and model evaluation, and performance was assessed using area under the curve (AUC) analysis. ResultsThe best-performing models achieved an AUC of 0.90 (95% CI: 0.88-0.93) for classifying indolent versus aggressive tumors and 0.76 (95% CI: 0.71-0.81) for benign versus malignant tumors. Models incorporating tumor size significantly improved classification accuracy. RF classifiers excelled in distinguishing indolent from aggressive tumors, while MLP classifiers performed better for benign versus malignant classification. ConclusionThe machine learning pipeline demonstrated high accuracy in differentiating aggressive from indolent renal tumors, offering valuable prognostic insights for personalized treatment. Tumor size was a critical factor, complementing CECT images and clinical data. These findings highlight the potential of ML techniques in enhancing renal tumor risk stratification.

Published in PLOS Digital Health · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.