Decoding Breast Cancer Heterogeneity via Multi-Omics Integration and Language Model-Based Interpretation
Yasrab, R.; Agrawal, R.; Saber-Ayad, M.; El-Hadidi, M.
Show abstract
We present a novel pipeline combining Multi-Omics Factor Analysis (MOFA) and fine-tuned Large Language Models (LLMs) to predict breast cancer subtypes using proteomics, DNA methylation, and RNA-Seq data. Breast cancer is a heterogeneous disease characterized by diverse molecular alterations across multiple biological layers, necessitating integrative approaches for accurate subtype classification. Our methodology leverages MOFA for dimensionality reduction to identify key latent factors driving het-erogeneity, followed by LLM fine-tuning on these multi-omics signatures to enhance prediction accuracy. MOFA analysis identified five key latent factors capturing distinct biological processes: immune response, cell cycle regulation, metabolic reprogramming, tumor microenvironment interactions, and DNA repair mechanisms. We extracted the top features per omics layer for each factor and performed Gene Set Enrichment Analysis (GSEA) to characterize their biological significance. Our LLM, trained on curated multi-omics signatures and clinical metadata encoded as structured text prompts, significantly outperformed conventional statistical models in subtype classification, achieving AUC=0.93 and accuracy=0.89, compared to Random Forest (AUC=0.87, accuracy=0.82) and SVM (AUC=0.85, accuracy=0.80). The superior performance of our approach is attributed to the LLMs ability to capture complex, non-linear relationships and hierarchical feature interactions across omics layers. This integrative pipeline provides both improved predictive performance and interpretable biological insights, offering potential for enhanced clinical decision-making in breast cancer management.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Revealing cancer driver genes through integrative transcriptomic and epigenomic analyses with Moonlight 95%
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 94%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 94%
Similar papers in this journal
- MethylSPWNet and MethylCapsNet: Biologically Motivated Organization of DNAm Neural Network, Inspired by Capsule Networks 92%
- Network- and Enrichment-based Inference of Phenotypes and Targets from large-scale Disease Maps 92%
- A personalised approach for identifying disease-relevant pathways in heterogeneous diseases 92%
Similar papers in this journal
- The significance of molecular heterogeneity in breast cancer batch correction and dataset integration 93%
- Automated quantification of Ki-67 expression in breast cancer from H&E-stained slides using a transformer-based regression model 93%
- Multiomic profiling of metastatic potential in estrogen receptor-positive human epidermal growth factor-negative breast cancer 92%
Similar papers in this journal
- Transformer-based deep learning integrates multi-omic data with cancer pathways 94%
- Autofluorescence Virtual Staining System for H&E Histology and Multiplex Immunofluorescence Applied to Immuno-Oncology Biomarkers in Lung Cancer 91%
- Evaluating the Radiation Sensitivity Index and 12-chemokine gene expression signature for clinical use in a CLIA laboratory 91%
Similar papers in this journal
- Co-Phosphorylation Networks Reveal Subtype-Specific Signaling Modules in Breast Cancer 95%
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 95%
- Deep Learning Approach to Identifying Breast Cancer Subtypes Using High-Dimensional Genomic Data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.