Field-of-view confounding shapes genetic discovery from self-supervised cardiac-imaging phenotypes
Pandey, D.; Narasimhan, V. M.
Show abstract
Self-supervised models increasingly convert medical images into quantitative phenotypes for biological discovery, but statistical reproducibility does not establish that a learned phenotype represents the intended anatomy. We trained a video masked-autoencoder on 69,932 UK Biobank cardiac cine-MRI studies and performed genome-wide association analysis of its latent representation. Although 18 of 20 leading axes were heritable with well-calibrated statistics, the representation encoded substantial field-of-view information: body size, stature and imaging centre (linear-probe R^2=0.55 for site); standard genomic-control and LD-score diagnostics did not identify this source of phenotype-level confounding. Restricting the field of view to the heart and residualising body and acquisition covariates before dimensionality reduction substantially attenuated linear and non-linear nuisance information while retaining cardiac signal. Adjusting the same covariates only during association testing attenuated nuisance associations but recovered substantially less of the cardiac-associated genetic signal, consistent with nuisance variation having already influenced the principal-component basis. The corrected representation identified new associated loci beyond those detected using supervised phenotypes at matched sample size, which shared genetic architecture selectively with cardiac-conduction traits and were localised to cardiac structures within the imaged field of view. Confounding in learned medical phenotypes can arise upstream of association testing, highlighting the importance of auditing and, where appropriate, correcting learned representations before association testing.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep Learning of Left Atrial Structure and Function Provides Link to Atrial Fibrillation Risk 94%
- Weakly supervised classification of rare aortic valve malformations using unlabeled cardiac MRI sequences 93%
- Computation and resource efficient genome-wide association analysis for large-scale imaging studies 93%
Similar papers in this journal
- Human visual cortex is organized along two genetically opposed hierarchical gradients with unique developmental and evolutionary origins 91%
- Cortical morphology at birth reflects spatio-temporal patterns of gene expression in the fetal human brain 90%
- Replicability of spatial gene expression atlas data from the adult mouse brain 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.