MOT: a Multi-Omics Transformer for multiclass classification tumour types predictions
OSSENI, M. A.; Tossou, P.; Laviolette, F.; Corbeil, J.
Show abstract
MotivationBreakthroughs in high-throughput technologies and machine learning methods have enabled the shift towards multi-omics modelling as the preferred means to understand the mechanisms underlying biological processes. Machine learning enables and improves complex disease prognosis in clinical settings. However, most multi-omic studies primarily use transcriptomics and epigenomics due to their over-representation in databases and their early technical maturity compared to others omics. For complex phenotypes and mechanisms, not leveraging all the omics despite their varying degree of availability can lead to a failure to understand the underlying biological mechanisms and leads to less robust classifications and predictions. ResultsWe proposed MOT (Multi-Omic Transformer), a deep learning based model using the transformer architecture, that discriminates complex phenotypes (herein cancer types) based on five omics data types: transcriptomics (mRNA and miRNA), epigenomics (DNA methylation), copy number variations (CNVs), and proteomics. This model achieves an F1-score of 98.37% among 33 tumour types on a test set without missing omics views and an F1-score of 96.74% on a test set with missing omics views. It also identifies the required omic type for the best prediction for each phenotype and therefore could guide clinical decisionmaking when acquiring data to confirm a diagnostic. The newly introduced model can integrate and analyze five or more omics data types even with missing omics views and can also identify the essential omics data for the tumour multiclass classification tasks. It confirms the importance of each omic view. Combined, omics views allow a better differentiation rate between most cancer diseases. Our study emphasized the importance of multi-omic data to obtain a better multiclass cancer classification. Availability and implementationMOT source code is available at https://github.com/dizam92/multiomic_predictions.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 95%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 95%
- Molecular Group and Correlation Guided Structural Learning for Multi-Phenotype Prediction 95%
Similar papers in this journal
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 96%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 95%
- Tensor decomposition- and principal component analysis-based unsupervised feature extraction to select more reasonable differentially expressed genes: Optimization of standard deviation versus state-of-art methods 94%
Similar papers in this journal
- MLcps: Machine Learning Cumulative Performance Score for classification problems 95%
- Imputing missing RNA-seq data from DNA methylation by using transfer learning based-deep neural network 95%
- Disease classification for whole blood DNA methylation: meta-analysis, missing values imputation, and XAI 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.