Back

MOSAIC: A Spectral Framework for Integrative Phenotypic Characterization Using Population-Level Single-Cell Multi-Omics

Lu, C.; Kluger, Y.; Ma, R.

2026-02-12 bioinformatics
10.64898/2026.02.10.705077 bioRxiv
Show abstract

Population-scale single-cell multi-omics offers unprecedented opportunities to link molecular variation to human health and disease. However, existing methods for single-cell multi-omics analysis are either cell-centric, prioritizing batch-corrected cell embeddings that neglect feature relationships, or feature-centric, imposing global feature representations that overlook inter-sample heterogeneity. To address these limitations, we present MOSAIC, a spectral framework that learns a high-resolution feature x sample joint embedding from population-scale single-cell multi-omics data. For each individual, MOSAIC constructs a sample-specific coupling matrix capturing complete intra- and cross-modality feature interactions, then projects these into a shared latent space via spectral decomposition. The joint feature x sample embedding defines each features connectivity profile per sample, enabling three downstream applications. Differential Connectivity analysis identifies features with regulatory network rewiring across conditions even when their abundance remains unchanged, revealing rewiring of proliferation programs in activated T cells from a vaccination cohort. Unsupervised subgroup detection isolates coherent feature modules to discover hidden patient subtypes, uncovering a stress-driven neuronal subtype within an HIV+ cohort. Clinical outcome prediction using connectivity-derived features complements abundance-based analysis, improving COVID-19 severity classification when integrated. MOSAIC provides a general-purpose framework for systems-level phenotypic characterization, bridging network-level discovery with clinical outcome prediction in population-scale single-cell studies.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.