Semi-supervised Omics Factor Analysis (SOFA) disentangles known sources of variation from latent factors in multi-omics data
Capraz, T.; Vöhringer, H. S.; Huber, W.
Show abstract
A fundamental design pattern in biomolecular studies is to assay the same set of samples (organisms, tissue biopsies, or individual cells) by multiple different omics assays. Group Factor Analysis (GFA) and its adaptation to high-dimensional settings, Multi-Omics Factor Analysis (MOFA), are widely used as a first-line approach to analyse such data and are effective in detecting patterns of correlation, organize them into so-called latent factors, and identify common and assay-specific factors. However, in many applications a subset of the found factors just rediscovers already known covariates (e.g., disease subtypes, environmental covariates) while others may represent genuine novelty. Here, we present Semi-supervised Omics Factor Analysis (SOFA), a method that incorporates known covariates into the model upfront and focuses the factor discovery on novel sources of variation. We show SOFAs effectiveness for discovering novel patterns by applying it to cancer, brain development and heart failure multi-omic data sets.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 97%
- Conserved epigenetic regulatory logic infers genes governing cell identity 97%
- Distinct gene programs underpinning 'disease tolerance' and 'resistance' in influenza virus infection 95%
Similar papers in this journal
Similar papers in this journal
- A tissue-aware machine learning framework enhances the mechanistic understanding and genetic diagnosis of Mendelian and rare diseases 96%
- PIFiA: Self-supervised Approach for Protein Functional Annotation from Single-Cell Imaging Data 96%
- hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments 96%
Similar papers in this journal
- Coralysis enables sensitive identification of imbalanced cell types and states in single-cell data via multi-level integration 97%
- Interpretable trajectory inference with single-cell Linear Adaptive Negative-binomial Expression (scLANE) testing 97%
- SPOTlight:Seeded NMF regression to Deconvolute Spatial Transcriptomics Spots with Single-Cell Transcriptomes 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.