Imaging Neuroscience
● MIT Press
All preprints, ranked by how well they match Imaging Neuroscience's content profile, based on 282 papers previously published here. The average preprint has a 0.19% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Kowalczyk, O. S.; Medina, S.; Venezia, A.; Tsivaka, D.; Ahmed, A. I.; Williams, S. C. R.; Brooks, J. C. W.; Lythgoe, D. J.; Howard, M. A.
Show abstract
Establishing the reliability of spinal cord functional magnetic resonance imaging (fMRI) is critical before employing it to assess experimental or clinical interventions. Previous studies have mapped human motor activity primarily to the ipsilateral ventral horn, aligning with myotomal and dermatomal projections. Despite these insights, the test-retest reliability of spinal fMRI remains under-investigated. Here we assessed spinal cord activation during a sensorimotor paradigm involving right-hand grasping and grip force estimation in 30 healthy volunteers. Participants completed two identical scanning visits, each time performing the same task twice, enabling the investigation of test-retest reliability both within a single experimental visit and between visits performed on different days. Aggregating all task runs, motor-evoked activation was observed in ipsilateral ventro-dorsal regions of spinal segmental levels C5-T1, as well as in medial regions of levels C2-C3. Despite highly reliable task performance (grip force) and fMRI signal quality (temporal signal-to-noise ratio), the reliability of motor activation was predominantly poor-to-fair both within and between visits, with notable variability in spatial distribution observed across task runs. Increasing the number of task runs per individual improved the robustness of group-level activation, as indexed by higher activated voxel count, larger cluster spatial extent, and attenuated t-statistic distribution. Although we demonstrated that motor-evoked activation corresponds to the known neuroanatomical organisation of motor circuits, its low test-retest reliability presents a challenge for wider applications of spinal fMRI. Understanding the drivers of low reliability in functional imaging is warranted, but we suggest that looking beyond measurement error is required, including careful consideration of inherent within-individual variability underpinned by neurophysiological and psychological factors.
Bhalerao, G. V.; Markiewicz, P.; Turnbull, J.; Thomas, D. L.; De Vita, E.; Parkes, L.; Thompson, G.; MacKewn, J.; Krokos, G.; Wimberley, C.; Hallett, W.; Su, L.; Malhotra, P.; Hoggard, N.; Taylor, J.-P.; Brooks, D.; Ritchie, C.; Wardlaw, J.; Matthews, P.; Aigbirho, F.; O'Brien, J.; Hammers, A.; Herholz, K.; Barkhof, F.; Miller, K.; Matthews, J.; Smith, S.; Griffanti, L.
Show abstract
Harmonisation is widely used to mitigate site- and scanner-related batch variability in multisite neuroimaging studies and is particularly critical in longitudinal clinical trials, where detection of subtle biological or treatment-related changes depends on reliable measurement across scanners and timepoints. However, the effectiveness of harmonisation in small, heterogeneous clinical datasets remains insufficiently understood, particularly in relation to subject-level variability and consistency across acquisition settings, and its impact on both removal of technical variability and preservation of biological variation in pooled multisite analyses. We systematically evaluated a range of image-based and statistical harmonisation methods using a clinically realistic multisite, multiscanner structural T1-weighted (T1w) MRI test-retest dataset comprising three controlled acquisition scenarios: repeatability, intra-scanner reproducibility and inter-scanner reproducibility. Methods were applied under different batch specifications (site, scanner, or both) and performance was assessed within each scenario and in pooled data using a multi-metric framework capturing both technical and biological variability in volumetric imaging-derived phenotypes (IDPs) relevant to aging and dementia research. Across IDPs, before harmonisation variability was lowest in the repeatability scenario (median variability=0.6 to 2.7%, rank consistency {rho} [≥]0.9), with modest increases under intra-scanner reproducibility (0.5 to 3.2%, {rho}=0.5 to 1.0) and substantially greater variability under inter-scanner reproducibility conditions (1.7 to 19.2%, {rho} =-0.1 to 0.9). These results offer important information to consider for multisite study design, including sample size calculation in clinical trials. Harmonisation performance was strongly context dependent, with clearer benefits emerged in inter-scanner scenarios where both variability reduction and improvements in subject-level consistency were observed. In pooled data, approaches that explicitly modelled site as batch and accounted for repeated-measure structure showed greater consistency across IDPs in batch effect mitigation and more accurately reflected underlying biological variation. Our evaluation metrics enabled disentangling the removal of global batch effect while highlighting residual variability at the phenotype-specific or multivariate levels. These findings demonstrate that harmonisation cannot be treated as a one-size-fits-all solution and must be interpreted relative to the acquisition context, dataset structure, and downstream analytic goals. Multi-metric evaluation under realistic clinical constraints is essential to support reliable and translatable neuroimaging inference by ensuring appropriate correction of batch effects while preserving longitudinal biological signals and sensitivity to clinically meaningful change in multisite studies.
Medina, M. C.; Reddy, N. A.; Bright, M. G.; Sitek, K. R.
Show abstract
Task-based precision mapping has become a promising technique in functional MRI (fMRI) to robustly characterize and map an individuals unique activity patterns. These experiments consist of acquiring extensive imaging data in one participant, ultimately improving the sensitivity and specificity of individual-specific functional localization. Despite its advantages, studies have primarily focused on understanding individual-specific cortical activation, preventing a holistic view of a systems-level functional response, and to date, best approaches for the statistical analysis of controlled task-based, densely sampled, whole-brain data have not yet been fully established. Therefore, in this study, we collected whole-brain (i.e. covering cortex, cerebellum, and brainstem) multi-echo densely sampled data of the auditory system, a system with major subcortical components, and evaluated activation sensitivity as well as activation stability across data subsets of commonly-used whole-brain and region-specific inference testing approaches. The whole-brain approaches involved standard voxel-level and cluster-level inference schemes with varying statistical thresholds and a non-parametric permutation inference approach. The region-specific approaches involved an exploratory top % t-statistics methods and non-parametric permutation inference approaches. We found that a whole-brain voxel-level approach with a false discovery rate (FDR) correction (p<0.05) presented highest sensitivity across regions and subjects as well as most consistent detection of expected auditory regions, even with lower scan duration. In addition, we found that a region-specific top % t-statistic approach may be a useful exploratory functional localization tool and a complementary method to standard inference testing approaches.
Kern, S.; Wittkuhn, L.; Buss, E.; Schuck, N.; Feld, G. B.
Show abstract
Studies in rodents and humans using invasive electrophysiology have established that neural replay is a ubiquitous phenomenon in the brain that is associated with a wide range of cognitive functions, including memory, planning and decision making. Yet, invasively recording in humans remains difficult, and hence knowledge about replay in humans remains scarce. Hence, to comprehensively understand replay in humans, we need reliable approaches that can detect it non-invasively. Several main non-invasive approaches have been proposed, but we lack a full comparative validation against known ground truth signals. In this study, we present FASTIMAGES, a benchmark dataset from seventy participants with parallel fMRI (n = 40, previously published) and MEG (n=30) recordings containing known neural sequences evoked by fast visual stimulation as well as functional localizer trials. The neural sequences were elicited by five different visual stimuli shown in sequences at speeds of 132, 164, 228 and 612 milliseconds onset-to-onset intervals. Using this dataset, we investigate two existing statistical methods for sequence detection, namely Temporally Delayed Linear Modelling (TDLM, developed for MEG by Liu et al., 2021) and Slope Order Dynamic Analysis (SODA, developed for fMRI by Wittkuhn & Schuck, 2021). We examine the underlying assumptions of each method, analyse their resulting strengths and weaknesses in application to MEG and fMRI. We demonstrate that both approaches excel in their native modality (TDLM for MEG and SODA for fMRI), with comparable effect sizes given idealized conditions in this benchmark. Cross-modality transfer remains challenging. Finally, the FASTIMAGES dataset provides data with known and clearly expressed sequences and can be used to benchmark and validate future sequence detection methods under idealized conditions.
Reddy, N. A.; Medina, M. C.; Northrop, J. N.; Zou, C.; Clements, R. G.; Nehrujee, A.; Sandhu, M.; Bright, M. G.
Show abstract
Multi-echo independent component analysis (ME-ICA) has been demonstrated to improve sensitivity and reliability of task functional magnetic resonance imaging (fMRI) data and, in particular, motor-task data with inherent task-correlated head motion. However, previous work has shown that an overly aggressive ME-ICA denoising approach may unintentionally remove task-related signal, while a more conservative approach may not effectively mitigate noise. While the effects of varied implementations of ME-ICA on signal and noise characteristics have been tested thoroughly in breath-hold data, the effects of similar modeling decisions have not been studied in motor-task data, which present with a more localized neural response. Here, we tested and compared the impacts of three analysis methods using rejected ME-ICA components as regressors in subject-level modeling: Aggressive (simple inclusion of ME-ICA regressors), Moderate (excluding task-correlated ME-ICA regressors from the model), and Conservative (orthogonalization of ME-ICA regressors to the core model and accepted ME-ICA components). We applied these methods to data from healthy and multiple sclerosis populations that included performance of hand-grasp, shoulder-abduction, and ankle-flexion tasks. We found that when the amount of head motion and its correlation with the task was high and the expected task-evoked signal was relatively low, the Conservative method led to significantly higher activation, t-statistics, and test-retest similarity in motor regions compared to the Aggressive method. Future motor-task studies may wish to implement similar models to prevent loss of motor signal, while still mitigating the effects of task-correlated head motion.
Li, M.; Wang, Y.; Jun, S.; Bringas Vega, M. L. L.; Valdes-Sosa, P. A.; An, L.; Jia, G.
Show abstract
Normative modeling expresses individual brain phenotypes as z-scores relative to a population norm, but in multi-site studies batch effects contaminate these z-scores and undermine their biomarker value. Existing approaches either harmonize data before fitting a normative model (ComBat+Normative), letting residual site effects leak into z-scores, or use parametric one-step methods (GAMLSS, HBR) that cannot flexibly model multivariate covariate interactions. We propose Generalized Normative Modeling (GNM), a onestep hierarchical framework that jointly estimates the global trajectory and site-specific effects via NUFFT-accelerated kernel regression with GCV bandwidth selection. Because the z-score is the ratio of batch-corrected residual to batch-corrected scale, residual site variance cancels algebraically -- a property we term self-correction. On ABIDE I cortical thickness (387 HC, 11 sites, 68 ROIs) and HarMNqEEG log-power spectra (1,564 subjects, 14 sites, 18 channels x 235 frequency bins), GNM produced the most site-invariant z-scores and best age-signal preservation among four methods. This work provides an open-source MATLAB toolbox with a declarative formula interface (https://github.com/LMNonlinear/Generalized-Normative-Modeling), enabling reliable individual-level inference in pooled multi-site cohorts and advancing the use of normative deviations as clinical biomarkers in precision psychiatry and neurology.
Thornberry, C.; Math, P.; Cohen Serra, M.; Seymour, R.; Nolan, C.; Whelan, R.
Show abstract
Optically pumped magnetometer magnetoencephalography (OPM-MEG) offers a wearable, movement-tolerant alternative to conventional cryogenic MEG, placing sensors closer to the scalp and, in principle, improving sensitivity to deep sources. This is advantageous for examining subcortical structures that are affected by ageing, disorders and disease, such as the hippocampus. However, it remains unclear whether well-established activity (such as the attenuation of theta oscillations during the imagination of novel scenes) can be recovered from the medial temporal lobe (MTL) with OPM-MEG, and whether an individual structural MRI is required. Here, fifteen adults completed a scene imagination task. Initially we applied a 12-parameter template warping coregistration pipeline to the full sample. Following source reconstruction, we recovered the expected attenuation of theta (4-8 Hz) power during scene imagination compared to a counting baseline, with a significant cluster of activity peaking in the left parahippocampal gyrus. The clusters centre of mass was localised to the left hippocampus (t = -3.55, p = 0.048, whole-brain FWE-corrected) and was mostly confined to the left medial temporal lobe. We further supported our findings by using an individual T1-weighted MRI reconstruction pipeline in six participants who had these scans available. The two approaches produced similar whole-brain topographies and localised the peak MTL theta effect to left hippocampus, with temporal-lobe conjunction centroids 3-mm apart. These findings provide evidence that the theta attenuation of the scene construction network can be recovered at the group level with OPM-MEG, without an individual MRI.
Lee, K. E.; Austin, T.; Uchitel, J.; Caballero Gaudes, C.; Collins-Jones, L. H.; Cooper, R.; Edwards, A.; Hebden, J.; Kromm, G.; Pammenter, K.; Blanco, B.
Show abstract
SignificanceDynamic functional connectivity (FC) in neonates is a growing area of interest due to the developmental significance of early functional networks. There are several emerging techniques to measure dynamic FC, adding new perspectives to well-studied static FC networks. Recent dynamic FC studies suggest that adult resting-state networks are driven by key moments of dynamic activity rather than sustained correlations. Co-activation pattern (CAP) analysis leverages this theory, clustering high-activity frames to identify recurring configurations of significant activity. High-density diffuse optical tomography (HD-DOT) is an infant-friendly modality that measures hemodynamic changes in the cortex and has been used to investigate static FC in term-aged infants. CAP analysis has not yet been applied to neonatal HD-DOT and may reveal temporal features of early functional networks that are not apparent from static methods. AimThis study applies CAP analysis to neonatal HD-DOT to characterize transient co-activation states and provide new insight into early functional brain networks beyond what is known from static FC methods alone. ApproachTask-free HD-DOT data were acquired from a cohort of sleeping term newborns at the Rosie Hospital, Cambridge UK (n = 44, postmenstrual age = 40+3 (range: 38+2-42+6) weeks). In each recording, the top 15% of seed-selected frames were clustered using the K-means algorithm for three regions of interest (ROIs: frontal, central, and parietal) to identify significant seed-associated patterns of co-activation or co-deactivation. These co-activation patterns (CAPs) were characterized for each infant using four metrics: consistency, fractional occupancy, dwell time, and transition likelihood. ResultsDistinct CAPs, reflecting the dynamic organization of neonatal cortical networks, were identified for frontal, central, and parietal regions. These CAPs showed high consistency scores, reflecting high intra-cluster spatial correlation and validating the efficacy of CAP analysis for newborn HD-DOT data. The CAP decomposition revealed significant patterns not observed in conventional static functional connectivity analyses. Several neonatal CAPs exhibited frontoparietal co-activation, potentially reflecting early default mode network activity, which is immature and modular for the first year after birth. ConclusionsThis work demonstrates the utility of CAP analysis with newborn HD-DOT and provides new insight into the dynamics of neonatal functional connectivity.
Hendriks, J.; Jansen, M. G.; Joules, R.; Pena-Nogales, O.; Elsen, F.; Povolotskaya, A.; Dijsselhof, M. B. J.; Rodrigues, P. R.; Barkhof, F.; Schrantee, A.; Mutsaerts, H.
Show abstract
Brain age is a promising biomarker for detecting atypical and pathological brain aging, but its accuracy and reliability depend critically on MRI quality. The impact of common MR image degradations such as motion, ghosting, blurring, and noise on brain age predictions remains unclear. In this study, we systematically assessed the effects of four simulated MRI artifact types, across ten severity levels, on brain age prediction using three widely used deep learning-based algorithms (Pyment, MIDI, MCCQR), in high-quality T1-weighted images of healthy adults (age range 18-85, 54% female). Artifact severity levels (1-10) were generated using a power-function mapping of TorchIO simulation parameters calibrated to the full PondrAI QC visual rating scale (from perfect to severely degraded image quality). Linear mixed-effects models with predicted brain age as dependent variable revealed a significant interaction between algorithm, artifact type, and severity (p<0.001), indicating algorithm-specific sensitivity to artifacts. In artifact-free scans, mean absolute error (MAE) was 4.6 years for MCCQR, 7.1 years for Pyment, and 9.1 years for MIDI. At severity level 10, MAE increased with up to 110% for Pyment, 112% for MCCQR, and 16% for MIDI (motion); and with up to 75% for Pyment, 135% for MCCQR, and 34% for MIDI (ghosting). Blurring had minimal impact at low-moderate levels, but at maximum severity MAE increased by 26% for Pyment and 137% for MCCQR, while MIDI remained largely stable. Noise minimally affected Pyment and MCCQR (MAE increases [≤]20%), but led to larger declines for MIDI (MAE increase 35%). The vulnerability of different algorithms highlights that training data, preprocessing strategies and underlying architectures influence robustness, emphasizing that artifact sensitivity is a key consideration when interpreting brain-age as a biomarker. Our results emphasize the need for artifact-aware evaluation and mitigation strategies when algorithms such as brain age are used in clinical research.
Li, M.; Wang, Y.; Shen, Y.; Jia, G.
Show abstract
Normative modeling quantifies individual deviation from population norms by estimating the conditional mean and variance of brain-derived measures as functions of clinically relevant parameters such as age. The rapid growth of multicenter consortia has created an urgent need for normative models that incorporate batch harmonization. Several harmonization methods based on linear mixed models--ComBat, GAMLSS, HBR, and Generalized Normative Modeling (GNM)--offer explicit formulations of the mean and variance, making them natural candidates for batch-harmonized normative modeling; yet the absence of a unified theoretical framework leaves it unclear whether and how these methods support the computation of batch-harmonized z-scores. We bridge this gap by writing existing harmonization methods as special cases of a single location-scale equation, y = m(x, {Theta})+{sigma}(x, {Theta}){varepsilon} , which we term the unified form of batch harmonization equation for normative modeling. The methods differ only in the functional forms of m and{sigma} , how batch parameters enter{Theta} , and how{Theta} is estimated. This unified form yields both harmonized data y* and site-invariant z-scores from the same model, providing a common theoretical language for harmonized normative modeling. Building on this framework, we evaluate the underlying regression engines (parametric, spline, Gaussian process, kernel, deep learning), sensitivity to outliers, computational scalability, and federated decomposability for privacy-preserving multi-center computation. By clarifying what each method assumes, what it delivers, and where the boundaries of current methodology lie, the unified equation establishes a principled foundation for method selection and charts a path toward reliable, scalable, and privacy-aware normative modeling across multi-center neuroimaging.
Avants, B.; Tustison, N. J.; Stone, J. R.
Show abstract
BackgroundMultimodal MRI (sMRI, dMRI, rsfMRI) encodes complementary aspects of brain structure and function; principled joint representations promise more sensitive and interpretable markers of brain health than single-modality features. MethodsWe evaluate Normative Neurological Health Embedding (NNHEmbed), a flexible multi-view framework that uses constrained cross-modal similarity objectives to learn low-dimensional embeddings. Models were trained and tested on the UK Biobank (n = 21,300) and evaluated for transfer and longitudinal sensitivity in independent cohorts (Normative Neuroimaging Library, Alzheimers Disease Neuroimaging Initiative, Parkinsons Progression Markers Initiative). ResultsNNHEmbed produced compact, biologically interpretable components that (a) maps to established neurocognitive systems (e.g., episodic memory, processing-speed, sensorimotor/basal-ganglia circuits), (b) generalizes across cohorts, and (c) captures within-subject change over time. Best configurations balance reconstruction fidelity and shared covariance, improving interpretability while preserving predictive utility. Case demonstrations illustrate individualized normative profiling across multiple visits. ConclusionsNNHEmbed yields stable, transferable multimodal embeddings suitable for normative mapping and longitudinal monitoring. Software, NNHEmbed configurations and derived bases are available for reproduction and reuse.
Tahayori, B.; Smith, R. E.; Vaughan, D. N.; Tailby, C.; Handwerker, D. A.; Pierre, E. Y.; Jackson, G. D.; Abbott, D. F.
Show abstract
Multi-echo functional Magnetic Resonance Imaging (fMRI) data are acquired by recording image volumes at multiple echo times and can be used to improve the separation of neural activity from noise. TE-Dependent ANAlysis (tedana) is an open-source software tailored to denoising of multi-echo fMRI data. The efficacy of denoising can however be inconsistent, often necessitating manual inspection that precludes its application in large-scale studies where processing is ideally fully automated. Here, we introduce Robusttedana, an optimised denoising pipeline that achieves adequate results at both single-subject and group level. Robust-tedana incorporates Marchenko-Pastur Principal Component Analysis (MPPCA) for effective thermal noise reduction, robust independent component analysis for stabilised signal decomposition, and a modified component classification process. We evaluated its performance on Multi-Band Multi-Echo (MBME) language-task fMRI data from the Australian Epilepsy Project (AEP) using objective measures, comparing to conventional fMRI analysis with and without multi-echo-based denoising. Experts manual evaluation was undertaken on a subset of these data to validate the objective measures. The proposed pipeline both mitigates the prevalence of erroneous attenuation of genuine task activation due to instability of single-subject analysis, and increases the magnitude of group-wise effects. Robust-tedana therefore facilitates advanced analysis of MBME fMRI data in an automated pipeline, including for clinical research assessment of individuals.
Barbarant, P.-L.; Meyniel, F.; Thirion, B.
Show abstract
Inter-individual variability poses a significant challenge in decoding brain activity across subjects. While standard anatomical registration procedures reduce morphological differences, they fail to capture functional variability between subjects. Functional alignment methods address this issue by establishing voxel-to-voxel correspondences between pairs of individuals, thereby constructing a shared functional space. This shared space can be extended at the group level by generating a functional template. However, despite the availability of toolboxes, functional templates remain underused in fMRI analysis. Adopting this approach is currently difficult due to the diversity of existing methods and the lack of clear guidelines. Comprehensive evaluations of functional templates are limited to movie-watching paradigms. Here, we extensively compare functional alignment methods (Optimal Transport, Procrustes, Ridge regression, and Shared Response Model) and template construction strategies (in-sample, out-of-sample, pairwise) within the more general framework of task decoding. In this framework, decoding accuracy measures how well individual activation patterns align. Across multiple tasks and datasets, we demonstrate that population templates built using Optimal Transport (a) yield the highest decoding accuracy, (b) are not significantly biased by individual subjects, which facilitates generalization to new subjects, and (c) preserve the cortical signal topography.
Horn, U.; Vannesjo, S. J.; Gross-Weege, N.; Trampel, R.; Revina, Y.; Kaptan, M.; Dabbagh, A.; Beghini, L.; Callot, V.; Todd, A.; Pine, K. J.; Moeller, H. E.; Finsterbusch, J.; Weiskopf, N.; Eippert, F.
Show abstract
Developments in functional magnetic resonance imaging (fMRI) at ultra-high field (UHF) now allow for insights into human brain function at mesoscopic scale. However, similar progress has not been achieved for the human spinal cord, despite its importance for reciprocal brain-body communication. Here, we therefore integrate substantial methodological improvements that enable previously unattainable high-resolution UHF spinal fMRI. Using a heat-pain paradigm in healthy volunteers, we investigate fMRI responses in the dorsal horn: we characterize their spatial pattern, establish their robustness and demonstrate that they can be reliably observed even for individual participants. Most importantly, we uncover two physiologically relevant spinal response components - phasic and tonic heat responses - that map differentially onto superficial and deep layers of the dorsal horn. Taken together, our novel UHF fMRI approach allows for differentiating responses across fundamental spinal processing units and will enable insights into human spinal cord function at mesoscopic scale in health and disease.
Navarro-Gonzalez, R.; Aja-Fernandez, S.; Planchuelo-Gomez, A.; de Luis-Garcia, R.
Show abstract
Foundation models (FMs) for brain magnetic resonance imaging (MRI) are increasingly adopted as pretrained backbones for clinical tasks such as brain age prediction, disease classification, and anomaly detection. However, if FM embeddings (internal representations) shift systematically across MRI scanners, downstream analyses built on them may reflect acquisition hardware rather than biology. No study has yet quantified this cross-scanner reproducibility. Here, we assess the cross-scanner reliability of brain MRI FM embeddings and investigate which design factors (pretraining strategy, network architecture, embedding dimensionality, and pretraining dataset scale) best explain the observed differences. Using the ON-Harmony travelling-heads dataset (20 participants, eight scanners, three vendors), we evaluate the embeddings of five architecturally diverse FMs and a FreeSurfer morphometric baseline via within- and between-scanner intraclass correlation coefficient (ICC), variance decomposition, and scanner fingerprinting. Reliability spanned the full spectrum: biology-guided models achieved good-to-excellent cross-scanner ICC (AnatCL: 0.970 [95\% confidence interval (CI): 0.94, 0.98]; y-Aware: 0.809 [0.63, 0.88]), matching or surpassing FreeSurfer (0.926 [0.83, 0.96]), whereas purely self-supervised models fell below the poor threshold (BrainIAC: 0.453, BrainSegFounder: 0.307, 3D-Neuro-SimCLR: 0.247), with 23--58\% of embedding variance attributable to scanner identity. The strongest correlate of cross-scanner reliability among the models evaluated was pretraining strategy: incorporating biological metadata (cortical morphometrics, age) into the contrastive objective produced scanner-robust embeddings, whereas architecture, dimensionality, and dataset scale did not predict reliability.
Kowalczyk, O. S.; Medina, S.; Tsivaka, D.; McMahon, S. B.; Williams, S. C. R.; Brooks, J. C. W.; Lythgoe, D. J.; Howard, M. A.
Show abstract
Resting fMRI studies have identified intrinsic spinal cord activity, which forms organised motor (ventral) and sensory (dorsal) resting-state networks. However, to facilitate the use of spinal fMRI in, for example, clinical studies, it is crucial to first assess the reliability of the method, particularly given the unique anatomical, physiological, and methodological challenges associated with acquiring the data. Here we demonstrate a novel implementation for acquiring BOLD-sensitive resting-state spinal fMRI, which was used to characterise functional connectivity relationships in the cervical cord and assess their test-retest reliability in 23 young healthy volunteers. Resting-state networks were estimated in two ways: (1) by extracting the mean timeseries from anatomically constrained seed masks and estimating voxelwise connectivity maps and (2) by calculating seed-to-seed correlations between extracted mean timeseries. Seed regions corresponded to the four grey matter horns (ventral/dorsal and left/right) of C5-C8 segmental levels. Test-retest reliability was assessed using the intraclass correlation coefficient (ICC) in the following ways: for each voxel in the cervical spine; each voxel within an activated cluster; the mean signal as a summary estimate within an activated cluster; and correlation strength in the seed-to-seed analysis. Spatial overlap of clusters derived from voxelwise analysis between sessions was examined using Dice coefficients. Following voxelwise analysis, we observed distinct unilateral dorsal and ventral organisation of cervical spinal resting-state networks that was largely confined in the rostro-caudal extent to each spinal segmental level, with more sparse connections observed between segments (Bonferroni corrected p < 0.003, threshold-free cluster enhancement with 5000 permutations). Additionally, strongest correlations were observed between within-segment ipsilateral dorso-ventral connections, followed by within-segment dorso-dorsal and ventro-ventral connections. Test-retest reliability of these networks was mixed. Reliability was poor when assessed on a voxelwise level, with more promising indications of reliability when examining the average signal within clusters. Reliability of correlation strength between seeds was highly variable, with highest reliability achieved in ipsilateral dorso-ventral and dorso-dorsal/ventro-ventral connectivity. However, the spatial overlap of networks between sessions was excellent. We demonstrate that while test-retest reliability of cervical spinal resting-state networks is mixed, their spatial extent is similar across sessions, suggesting that these networks are characterised by a consistent spatial representation over time.
Bedard, S.; Kaptan, M.; Indriolo, T.; Law, C. S.; Pfyffer, D.; Lee, L.; Ratliff, J.; Hu, S.; Tharin, S.; Smith, Z. A.; Glover, G. H.; Mackey, S.; Cohen-Adad, J.; Weber, K. A.
Show abstract
Sensory organization at the spinal segment level is commonly inferred from dermatomal maps that assume a fixed correspondence between cutaneous regions and spinal segments. However, based on the complexities of spinal neuroanatomy and neurophysiology, the distribution of sensory signals within the cord may be broader and less segment-specific than dermatomal maps suggest, leaving the segment-level localization of sensory-evoked activity in humans uncertain. Spinal cord functional magnetic resonance imaging (fMRI) is currently the only technique capable of noninvasively mapping sensory activity with high spatial resolution in the human spinal cord. However, its application remains technically challenging and is limited by the uncertainty in segmental localization. In this study, we leveraged recent advancements in spinal cord fMRI, including spinal nerve rootlet-based spatial normalization, to investigate how sensory information is represented and distributed within the human spinal cord during electrocutaneous stimulation of the third digit of the right hand (i.e., C7 dermatome). Forty healthy adults were scanned with electrocutaneous stimulation at four individualized intensities across multiple runs to quantify (i) the rostrocaudal distribution of sensory-evoked activity, (ii) intensity-dependent changes in detectability and localization, and (iii) the effect of normalization strategy on segmental localization. Across participants, stimulation produced activation localized in the lower cervical cord (e.g., C6-C8), with the most consistent segmental localization near C7. Stronger stimulation increased detectability and produced more consistent segmental localization across participants. Importantly, normalization that incorporated nerve rootlet landmarks sharpened localization and improved sensitivity relative to conventional intervertebral disc-based alignment. This highlights the value of functionally relevant anatomical landmarks for group inference in the spinal cord. Responses were strongest in the initial run and attenuated with repetition, suggesting habituation or adaptation that can bias multi-run paradigms if unmodeled. Together, our results define practical acquisition and analysis conditions (e.g., stimulation strength, anatomical alignment strategy, and run structure) under which segment-level spinal sensory responses can be detected, thereby supporting more reliable studies of human spinal cord future basic and translational studies, including pain mechanisms, sensory function, and spinal injury.
Choi, S.; Shaw, J.; Cooper, R.; Corcoran, M.; Sathe, S.; Hayes, R.; Elder, I.; Lucas, A.; Vadali, C.; Stein, J.; Jalbrzikowski, M.
Show abstract
Portable low-field MRI systems are a promising complement to conventional high-field systems, enabling broader access to MRI. However, correspondence in cortical thickness estimates between low- and high-field MRI in young people remains limited despite its importance for neurodevelopment and psychopathology. To evaluate how multiple low-field image processing approaches improve cortical thickness correspondence with high-field MRI in a large sample of young individuals, we collected ultra-low-field (64mT) and high-field (3T) MRI data from a community sample of young people. We applied deep learning-based image processing approaches (SynthSR v1.0, SynthSR v2.0, recon-all-clinical, and recon-any) to low-field data acquired across multiple sequences (T1- and T2-weighted) and orientations (axial, coronal, sagittal, and multi-orientation), with and without resampling and/or co-registration. We assessed global, lobar, and regional cortical thickness correspondence with 3T MRI measures using Pearson and intraclass correlations. We compared pipelines using Steigers Z-tests and Fishers Z-tests. A total of 150 individuals (mean age, 18.63{+/-}5.07; 80 female) were included. We observed the highest global correspondence with recon-all-clinical applied to coronal T1-weighted images (r=0.40, pFDR=2.6e-05). At the lobar and regional levels, multi-orientation T2-weighted images processed with recon-all-clinical showed the highest correspondence across the greatest number of regions (4/12 lobes; 13/68 regions). The highest correspondence and largest improvements were in frontal, cingulate, and temporal regions, including the right pars triangularis (r=0.52, pFDR=4.78e-11; Z=4.78, pFDR=4.25e-06), right caudal anterior cingulate (r=0.47, pFDR=3.83e-09; Z=5.46, pFDR=1.32e-07), and left parahippocampal (r=0.58, pFDR=2.98e-14; Z=5.17, pFDR=6.01e-07). We observed significantly improved cortical thickness correspondence in low-field MRI in young people. The recon-all-clinical pipeline yielded moderate correspondence, particularly in frontal, cingulate, and temporal regions. Our results highlight the potential of low-field MRI as an affordable and scalable approach for assessing cortical thickness in young people.
Turnbull, J.; Bhalerao, G.; Dawson, R.; Lange, F.; Alfaro-Almagro, F.; Smith, S.; Griffanti, L.
Show abstract
Big neuroimaging data enable researchers to study subtle structural and functional brain changes and relationships between brain characteristics and genetics, lifestyle, and disease factors. However, substantial effort is needed to minimise technical, non-biological differences between data batches to avoid incorrect inferences. In this study, we address a previously identified bias in UK Biobank FreeSurfer IDPs derived from only the T1 image compared to those using both T1 and T2-FLAIR by treating the bias as a batch effect and using harmonisation approaches. We investigate and characterise this bias through direct within-participant comparison at the image and IDP level, comparing the results with those seen in the wider UKB sample. We then assess different methods of addressing the effect of missing T2-FLAIR, starting from simple linear regression before moving to ComBat, a widely used harmonisation method, testing different approaches for applying ComBat and showing its similarity to simple linear regression. Finally, we examine how ComBat estimates vary with batch and sample size. Our results show clear benefits in using both T1 and T2-FLAIR data in FreeSurfer, as opposed to just the T1, which is more common, with the pial surface fitting being less likely to fail and showing greater biologically plausible inter-subject variability. This is particularly important for cortical thickness IDPs, where T2-FLAIR omission leads to reduced true variability and systematic underestimation, as shown through within-participant repeat testing. We demonstrate that ComBat can address this bias, with its standard use (i.e., applied separately on different IDP categories) showing the best improvement in cortical thickness measures where the bias is strongest, and we find that it is important not to pool ComBat priors across different classes of IDPs. Our proposed version of ComBat with a reference batch (i.e., estimating mean and variance only from data with T2-FLAIR available) performed best in recovering both mean and variance differences between batches across different IDP classes and offers a promising approach for cases where a reference batch is clearly identifiable. While ComBat reliably corrects mean (additive) batch effects with relatively small sample sizes ({approx}30 subjects per batch), we show that its variance (multiplicative) correction is substantially less stable, requiring much larger sample sizes and becoming unreliable when batches are small or imbalanced, or when there is a large variance difference between them.
Okell, T. W.; Xu, X.; Craig, M.; Alfaro-Almagro, F.; Thomas, D. L.; De Vita, E.; Garratt, S.; Nichols, T. E.; Guenther, M.; Matthews, P. M.; Miller, K. L.; Smith, S. M.; Chappell, M. A.
Show abstract
Blood flow to the brain is a sensitive marker of neuronal activity as well as of a number of diseases, including stroke, tumours and neurodegenerative conditions. Arterial spin labelling (ASL) is a non-invasive magnetic resonance imaging (MRI) method that can map brain perfusion, but the ability to identify relationships between blood flow and lifestyle, genetics and disease has been limited by the scale of ASL studies to date. Here, we describe the inclusion of ASL in the repeat-imaging component of the UK Biobank imaging study, a prospective epidemiological study that has acquired 100,000 first-scan datasets and aims to accumulate over 60,000 repeat-scan datasets in predominantly healthy participants, along with rich information about lifestyle factors, genetics and long-term health outcomes. The imaging protocol and analysis pipeline are outlined, along with preliminary analyses of the first 7,157 subjects (more than twice as many as the largest previous ASL study). Significant associations with a range of factors are found, including those relating to the heart and blood vessels, alcohol consumption, cognitive tasks, white matter lesions and health information, such as hearing loss and depression. ASL is shown to be more sensitive to many of these factors than other imaging modalities, complementing the existing range of structural and functional measures available in the protocol. This resource is available to researchers worldwide, which we hope will facilitate new insights into healthy brain function and pathophysiology, and potentially allow the identification of early markers of disease as long-term health outcomes accumulate.