MOFA-FLEX: A Factor Model Framework for Integrating Omics Data with Prior Knowledge
Qoku, A.; Rohbeck, M.; Walter, F. C.; Kats, I.; Stegle, O.; Buettner, F.
Show abstract
Latent factor models are first-line analysis approaches for single- and multi-omics data, essential for data integration, alignment, and biological signal discovery. To cater for new technologies and experimental designs, bespoke extensions of factor models have been proposed, incorporating spatial structure, temporal dynamics and the noise characteristics of single-cell assays. However, the development of tailored methods and software for individual use cases is laborious and requires advanced statistical and domain expertise, posing a significant barrier to users. To address this, we here propose MOFA-FLEX, a flexible and modular factor analysis framework designed for customisable modelling across diverse multi-omics data scenarios. Built on probabilistic programming, MOFA-FLEX unifies previously isolated extensions of factor analysis - including flexible priors, non-negativity constraints, supervision signals, and alternative data likelihoods - allowing models to be configured declaratively without requiring manual engineering. Additionally, MOFA-FLEX features a novel domain knowledge module to inform and connect latent factors to gene programs. We demonstrate MOFA-FLEX across multiple applications, showing (i) improved robustness in recovering gene programs from noisy prior knowledge in scRNA-seq data; (ii) effective disentanglement of technical and biological variation in multi-omic CITE-seq; and (iii) tailored spatial modelling that reveals spatially organised disease-associated gene programs in breast cancer.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 96%
- scTrace+: enhance the cell fate inference by integrating the lineage-tracing and multi-faceted transcriptomic similarity information 95%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.