Back

Uncovering Multi-Omics Profiles of Population Heterogeneity: A Cluster-Based Bayesian Approach

Shi, R.; Xiong, Z.; Banaschewski, T.; Barker, G. J.; Bokde, A. L. W.; Flor, H.; Garavan, H.; Gowland, P.; Grigis, A.; Heinz, A.; Martinot, J.-L.; Martinot, M.-L. P.; Artiges, E.; Nees, F.; Orfanos, D. P.; Poustka, L.; Smolka, M. N.; Hohmann, S.; Vaidya, N.; Walter, H.; Whelan, R.; Schumann, G.; Desrivieres, S.; Lin, X.; Feng, J.

2025-12-20 genetic and genomic medicine
10.64898/2025.12.19.25342631 medRxiv
Show abstract

Understanding the heterogeneous nature of genetic effects is critical for advancing our knowledge of the genetic architecture of complex traits and developing personalized management strategies. However, existing methods often rely on pre-specified modifying variables to model this heterogeneity, limiting their ability to capture effects driven by complex or unobserved factors. Here, we propose MOCHA (Multi-Omics Clustering for Heterogeneous Association), a novel Bayesian analytical paradigm that identifies latent population subgroups with distinct genetic effects directly from multi-omics data, without requiring a priori variable specification. Extensive simulations confirm that MOCHA accurately identifies the underlying clustering structure, demonstrates superior performance in identifying and ranking features with cluster-specific effects, and provides reliable parameter estimates. Applying MOCHA to genomic and transcriptomic data from the IMAGEN study, we identified two distinct neurodevelopmental clusters associated with adolescent inhibitory control. Post-hoc characterization of these clusters provided novel insights into the mechanisms of brain plasticity, demonstrating the methods practical utility and interpretability.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.