Back

Identification of Persistent Radiomics Feature Co-occurrence Across Diverse Tissue Types and Individuals: A Network-Based Analysis of the RADAPT CT Atlas

Amiri, S.; Afshar, P.; Rohban, M. H.

2026-07-19 radiology and imaging
10.64898/2026.07.17.26358252 medRxiv
Show abstract

Objectives. Radiomics pipelines extract hundreds of quantitative features that are widely known to be redundant, but the structure of this redundancy is usually treated as a per-dataset nuisance to be pruned away. We tested the alternative hypothesis that a substantial number of feature-feature correlations are universal: they persist across patients and across anatomically distinct structures because they reflect shared mathematical and image-statistical properties of how the image is summarised, rather than properties of the tissue being imaged. Materials and Methods. We re-analysed the publicly available Radiomics Atlas Dataset of normal Abdominal and Pelvic CT (RADAPT), restricting the analysis to the 526 non-contrast-enhanced examinations of the 531-subject atlas and to the 107 original (non-filtered) PyRadiomics features. The 53 segmented structures were grouped into four broad anatomical categories -- bones, muscles, vessels, and parenchymal organs. RADAPT is distributed as one Excel file per structure, with patients as rows and features as columns. Within each structure file we z-score-normalised every feature across patients, computed the absolute Spearman correlation matrix, and retained edges with |{rho}| [≥] {tau} for {tau} in {0.70, 0.80, 0.90}. We then intersected the edge sets across all structure files to obtain a "universal" correlation graph, in which an edge survives only if it exceeds the threshold in every structure (each estimated across the full patient sample). Stable feature communities were defined as the maximal cliques of this graph. Robustness to patient sampling was tested by repeating the entire pipeline on five independent random splits of each file into two patient halves (10 sub-cohorts per threshold), and the implementation was independently reproduced in R. Results. Despite the strictness of the global-intersection criterion, 34, 24, and 14 stable feature communities survived at {tau} = 0.70, 0.80, and 0.90 respectively, with the largest cliques containing six members at {tau} = 0.70 and {tau} = 0.80 and five members at {tau} = 0.90. The community structure was clearly interpretable: separate cliques captured (i) variance-like intensity dispersion, (ii) long-run / low-frequency (coarse) texture, (iii) high gray-level texture, (iv) low gray-level texture, (v) volume and surface shape, and (vi) local-homogeneity and energy/entropy duals. On random-half resampling the exact-match recovery rate of these communities was 81.5 %, 86.7 %, and 80.7 % across the three thresholds; departures from exact recovery were almost always a single boundary feature added or dropped, consistent with finite-sample fluctuation of near-threshold edges rather than structural instability. The R re-implementation reproduced the Python results exactly. Conclusion. A substantial portion of radiomics feature collinearity is universal across patients and tissues. We distinguish two layers within it: trivial near-algebraic duals that are universal by construction, and non-trivial cross-matrix-family communities that are the genuine empirical finding. Together they provide an interpretable, definition-grounded basis for aggressive dimensionality reduction, for retrospectively reconciling apparently different feature selections in the literature, and for moving radiomics pipelines toward organ-agnostic, more reproducible models. Clinical relevance statement. Selecting a single representative feature from each universal community shrinks the original-feature space by roughly an order of magnitude without sacrificing biologically distinct information. For example, the five variance-family members (first-order Variance, GLCM SumSquares, GLCM ClusterTendency, GLDM and GLRLM GrayLevelVariance) can be replaced by a single representative, removing redundant degrees of freedom that would otherwise inflate model variance; and labelling each retained feature by its community lets two studies that selected different variance-family names be recognised as having found the same signal, simplifying model development and improving cross-cohort generalisability in clinical CT workflows.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 12%
14.8%
2
Communications Biology
993 papers in training set
Top 0.5%
7.7%
3
Science Advances
1243 papers in training set
Top 3%
6.6%
4
Imaging Neuroscience
282 papers in training set
Top 0.9%
6.2%
5
Scientific Reports
3612 papers in training set
Top 13%
6.2%
6
GigaScience
212 papers in training set
Top 0.6%
5.4%
7
Medical Image Analysis
35 papers in training set
Top 0.2%
5.1%
50% of probability mass above
8
Human Brain Mapping
329 papers in training set
Top 1%
4.8%
9
PLOS Computational Biology
1863 papers in training set
Top 8%
4.8%
10
Medical Physics
14 papers in training set
Top 0.2%
3.4%
11
NeuroImage
903 papers in training set
Top 3%
3.4%
12
eLife
5828 papers in training set
Top 36%
3.1%
13
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 25%
1.9%
14
PLOS ONE
5266 papers in training set
Top 47%
1.9%
15
Communications Medicine
113 papers in training set
Top 2%
1.7%
16
Advanced Science
286 papers in training set
Top 6%
1.3%
17
Journal of Medical Imaging
11 papers in training set
Top 0.2%
1.3%
18
European Radiology
15 papers in training set
Top 0.5%
1.1%
19
Nature Machine Intelligence
70 papers in training set
Top 2%
1.1%
20
Science Translational Medicine
127 papers in training set
Top 3%
0.9%
21
Aperture Neuro
20 papers in training set
Top 0.4%
0.9%
22
Nature Medicine
125 papers in training set
Top 3%
0.9%
23
Photoacoustics
12 papers in training set
Top 0.3%
0.8%
24
npj Digital Medicine
118 papers in training set
Top 3%
0.8%
25
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
0.8%
26
npj Precision Oncology
53 papers in training set
Top 2%
0.6%
27
Scientific Data
209 papers in training set
Top 3%
0.6%
28
PLOS Biology
486 papers in training set
Top 15%
0.6%