Back

Machine learning for integrative multi-omics clustering and feature gene identification

Zhang, Y.; Liu, L.; Ma, L.; Zhang, Z.

2025-12-02 bioinformatics
10.64898/2025.11.28.691099 bioRxiv
Show abstract

Multi-omics integrative analysis is pivotal for elucidating complex molecular mechanisms and biological processes, yet remains challenging due to the high dimensionality and heterogeneity of multi-omics data. Here we describe MIA, a machine learning framework for multi-omics integrative analysis. Unlike existing algorithms that rely on two-dimensional representations, MIA employs a three-dimensional tensor representation coupled with tensor decomposition, Fuzzy C-Means, and an enhanced random forest model to jointly enable accurate sample stratification and feature discovery. Benchmarking on simulated datasets demonstrates that MIA achieves higher accuracy in both clustering and feature identification by comparison with extant algorithms. Application to TCGA datasets further shows the ability of MIA to stratify samples and identify features associated with significantly clinical outcomes. Notably, as applied to glioblastoma, MIA is capable to detect three previously uncharacterized subtypes with distinct prognostic profiles and uncover critical feature genes linked to glioblastoma subtyping and therapeutic response. Collectively, these results establish MIA as a generalizable computational framework for multi-omics integrative analysis, enabling systematic molecular subtyping and significant feature discovery across complex biological systems.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.