Back

Multivariate Random Forests for Cross-Modal Multi-Omics Integration

Zhang, W.; Wang, L.; Franzmann, E. J.; Chen, X. S.

2026-06-22 bioinformatics
10.64898/2026.06.17.732933 bioRxiv
Show abstract

Multi-omics studies are widely used across many areas of biomedical research. In many diseases, some signals are shared across data types, while others are strongest in a single omics layer. Current multi-omics clustering methods often either merge all data types into a single representation, which can blur biology that is strong in one layer, or rely on linear structure that may miss more complex relationships across data types. We introduce O_SCPLOWMULTIC_SCPLOWRF, a random-forest-based method that handles complex data types and separates shared and modality-specific structure for multi-omics data. O_SCPLOWMULTIC_SCPLOWRF learns sample similarities across omics layers from multivariate random forests, combines them across data types, and uses the resulting weights to estimate the part of each omics layer that is predictable from the others. The remaining residual is treated as modality-specific signal, allowing shared and modality-specific similarities to be clustered separately. In simulations, O_SCPLOWMULTIC_SCPLOWRF recovered shared clusters as well as or better than established integrative methods while more reliably separating modality-specific signal under nonlinear data structures. In TCGA head and neck squamous cell carcinoma, the shared component aligned with the main subtype structure across established reference classifications, while gene- and miRNA-specific components revealed additional immune and developmental biology. In the ADNI cohort with matched blood DNA methylation and structural MRI, the shared cross-modal aging signal was associated with future conversion to mild cognitive impairment or Alzheimers disease, and a DNAm-specific residual signal showed exploratory additional information. These results show that O_SCPLOWMULTIC_SCPLOWRF can recover a common disease axis while retaining biologically meaningful signals specific to one data type. O_SCPLOWMULTIC_SCPLOWRF is available as an open-source R package at https://github.com/novawz/multiRF.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.