Quantifying the advantage of multimodal data fusion for survival prediction in cancer patients
Nikolaou, N.; Salazar, D.; RaviPrakash, H.; Goncalves, M.; Mulla, R.; Burlutskiy, N.; Markuzon, N.; Jacob, E.
Show abstract
The last decade has seen an unprecedented advance in technologies at the level of high-throughput molecular assays and image capturing and analysis, as well as clinical phenotyping and digitization of patient data. For decades, genotyping (identification of genomic alterations), the casual anchor in biological processes, has been an essential component in interrogating disease progression and a guiding step in clinical decision making. Indeed, survival rates in patients tested with next-generation sequencing have been found to be significantly higher in those who received a genome-guided therapy than in those who did not. Nevertheless, DNA is only a small part of the complex pathophysiology of cancer development and progression. To assess a more complete picture, researchers have been using data taken from multiple modalities, such as transcripts, proteins, metabolites, and epigenetic factors, that are routinely captured for many patients. Multimodal machine learning offers the potential to leverage information across different bioinformatics modalities to improve predictions of patient outcome. Identifying a multiomics data fusion strategy that clearly demonstrates an improved performance over unimodal approaches is challenging, primarily due to increased dimensionality and other factors, such as small sample sizes and the sparsity and heterogeneity of data. Here we present a flexible pipeline for systematically exploring and comparing multiple multimodal fusion strategies. Using multiple independent data sets from The Cancer Genome Atlas, we developed a late fusion strategy that consistently outperformed unimodal models, clearly demonstrating the advantage of a multimodal fusion model.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 93%
- A Deep Survival EWAS approach estimating risk profile based on pre-diagnostic DNA methylation: an application to Breast Cancer time to diagnosis 93%
- Survival analysis of pathway activity as a prognostic determinant in breast cancer 93%
Similar papers in this journal
- Empirical methods for the validation of Time-To-Event mathematical models taking into account uncertainty and variability: Application to EGFR+ Lung Adenocarcinoma. 93%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 92%
- Ranking Cancer Drivers via Betweenness-based Outlier Detection and Random Walks 92%
Similar papers in this journal
Similar papers in this journal
- Survival Prediction Landscape: An In-Depth Systematic Literature Review on Activities, Methods, Tools, Diseases, and Databases 95%
- Data-driven Discovery of Mathematical and Physical Relations in Oncology Data using Human-understandable Machine Learning 92%
- Discovering early imaging biomarkers of osteoradionecrosis in oropharyngeal cancer by characterization of temporal changes in computed tomography mandibular radiomic features 91%
Similar papers in this journal
- Two-step multi-omics modelling of drug sensitivity in cancer cell lines to identify driving mechanisms 93%
- MFmap: A semi-supervised generative model matching cell lines to tumours and cancer subtypes 93%
- Assessing reliability of intra-tumor heterogeneity estimates from single sample whole exome sequencing data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.