Back

Large-scale population neuroimaging reveals latent subgroup structure in functional brain organisation

Farahibozorg, S. R.; Smith, S. M.; Elliott, L. T.; Woolrich, M. W.

2026-07-19 neuroscience
10.64898/2026.07.17.739129 bioRxiv
Show abstract

Large-scale functional MRI datasets provide resources to understand inter-individual variation in human brain function and relate this variation to behaviour and health. However, most existing approaches fail to bridge the gap between population-average and individual-specific modelling, limiting the identification of structured subgroup heterogeneity across individuals. Here we develop a scalable framework for unsupervised subgroup discovery in population-scale resting-state fMRI data from 19,993 UK Biobank participants. Using stochastic Probabilistic Functional Modes, we estimate population-informed individualised spatial topographies of resting-state networks and derive high-dimensional functional fingerprints for each participant. We then identify latent subgroups by applying Gaussian mixture modelling independently to each fingerprint feature, yielding hundreds of reproducible subgroup definitions across 1,000 functional dimensions. We report approximately 5,700 significant differences between subgroups in a range of non-imaging phenotypes related to cognition, lifestyle, physical and mental health. Spatial organisation of the brain networks reveals distinct subgroup differences in sensory-motor and higher-order cognitive systems, in addition to correspondence with regional patterns of genetic variability across the brain. Together, these results demonstrate that large-scale functional neuroimaging contains rich latent subgroup structure linked to behavioural and biological variation. Our framework provides an interpretable and scalable basis for stratified models of human brain function and population neuroscience.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.