Back

Heterogeneity, Longitudinal Decline, and Metabolic Risk in MRI-Based Quantification of 20 Individual Hip and Thigh Muscles

Whitcher, B.; Raza, H.; Basty, N.; Thanaj, M.; Bell-Bradford, C.; Niglas, M.; Bell, J. D.; Thomas, E. L.; Amiras, D.

2026-02-27 radiology and imaging
10.64898/2026.02.25.26347009 medRxiv
Show abstract

Quantifying muscle health at scale has been limited by the difficulty of segmenting individual muscles on MRI. We developed an automated 3D deep-learning framework that segments 20 bilateral hip and thigh muscles from Dixon MRI, enabling muscle level quantification of volume and relative fat fraction (rFF). Applied to 10,840 baseline and 2,766 longitudinal UK Biobank scans, this framework supports population-scale phenotyping across demographic, metabolic and treatment exposures. Segmentation accuracy was robust, and increased with muscle size. Men had greater muscle volumes, whereas women showed consistently higher rFF. Fat infiltration was highest in postural and pelvic-stabilising muscles and lowest in the quadriceps, revealing pronounced anatomical heterogeneity. Over two years, most muscles showed small but consistent volume declines, with losses more uniform in men and more heterogeneous in women; rFF increased more prominently in women, suggesting early compositional deterioration. In T2D, men showed widespread volume loss and elevated rFF, whereas women showed minimal volume loss and heterogeneous fat changes, revealing sex-specific disease signatures. Automated muscle-specific MRI phenotyping resolves structural and compositional changes obscured by compartment-level measures and provides a scalable platform for population-level studies of musculoskeletal ageing, metabolic disease, and therapeutic response.

Published in Scientific Reports (predicted rank #2) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.