Machine learning for integrative multi-omics clustering and feature gene identification
Zhang, Y.; Liu, L.; Ma, L.; Zhang, Z.
Show abstract
Multi-omics integrative analysis is pivotal for elucidating complex molecular mechanisms and biological processes, yet remains challenging due to the high dimensionality and heterogeneity of multi-omics data. Here we describe MIA, a machine learning framework for multi-omics integrative analysis. Unlike existing algorithms that rely on two-dimensional representations, MIA employs a three-dimensional tensor representation coupled with tensor decomposition, Fuzzy C-Means, and an enhanced random forest model to jointly enable accurate sample stratification and feature discovery. Benchmarking on simulated datasets demonstrates that MIA achieves higher accuracy in both clustering and feature identification by comparison with extant algorithms. Application to TCGA datasets further shows the ability of MIA to stratify samples and identify features associated with significantly clinical outcomes. Notably, as applied to glioblastoma, MIA is capable to detect three previously uncharacterized subtypes with distinct prognostic profiles and uncover critical feature genes linked to glioblastoma subtyping and therapeutic response. Collectively, these results establish MIA as a generalizable computational framework for multi-omics integrative analysis, enabling systematic molecular subtyping and significant feature discovery across complex biological systems.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TrimNN: Characterizing cellular community motifs for studying multicellular topological organization in complex tissues 97%
- Triple-effect correction for Cell Painting data with contrastive and domain-adversarial learning 96%
- HEARTSVG: a fast and accurate method for spatially variable gene identification in large-scale spatial transcriptomic data 96%
Similar papers in this journal
- Predicting MammaPrint Recurrence Risk from Breast Cancer Pathological Images Using a Weakly Supervised Transformer 96%
- DeDoc2 identifies and characterizes the hierarchy and dynamics of chromatin TAD-like domains in the single cells 95%
- Cancer-like fragmentomic characteristics of somatic variants in cell-free DNA 94%
Similar papers in this journal
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 96%
- CosGeneGate Selects Multi-functional and Credible Biomarkers for Single-cell Analysis 96%
- HyGAnno: Hybrid graph neural network-based cell type annotation for single-cell ATAC sequencing data 95%
Similar papers in this journal
- Gene Function Revealed at the Moment of Stochastic Gene Silencing 95%
- Manifold learning analysis suggests novel strategies for aligning single-cell multi-modalities and revealing functional genomics for neuronal electrophysiology 95%
- Chrysalis: decoding tissue compartments in spatial transcriptomics with archetypal analysis 94%
Similar papers in this journal
- stGCL: A versatile cross-modality fusion method based on multi-modal graph contrastive learning for spatial transcriptomics 96%
- SOAPy: a Python package to dissect spatial architecture, dynamics and communication 96%
- GeneWalk identifies relevant gene functions for a biological context using network representation learning 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.