Benchmarking of analytical combinations for COVID-19 outcome prediction using single-cell RNA sequencing data
Cao, Y.; Ghazanfar, S.; Yang, P.; Yang, J.
Show abstract
The advances of single-cell transcriptomic technologies have led to increasing use of single-cell RNA sequencing (scRNA-seq) data in large-scale patient cohort studies. The resulting high-dimensional data can be summarised and incorporated into patient outcome prediction models in several ways, however, there is a pressing need to understand the impact of analytical decisions on such model quality. In this study, we evaluate the impact of analytical choices on model choices, ensemble learning strategies and integration approaches on patient outcome prediction using five scRNA-seq COVID-19 datasets. First, we examine the difference in performance between using each single-view feature space versus multi-view feature space. Next, we survey multiple learning platforms from classical machine learning to modern deep learning methods. Lastly, we compare different integration approaches when combining datasets is necessary. Through benchmarking such analytical combinations, our study highlights the power of ensemble learning, consistency among different learning methods and robustness to dataset normalisation when using multiple datasets as the model input. Summary key pointsO_LIThis work assesses and compares the performance of three categories of workflow consisting of 350 analytical combinations for outcome prediction using multi-sample, multi-conditions single-cell studies. C_LIO_LIWe observed that using ensemble of feature types performs better than using individual feature type C_LIO_LIWe found that in the current data, all learning approaches including deep learning exhibit similar predictive performance. When combining multiple datasets as the input, our study found that integrating multiple datasets at the cell level performs similarly to simply concatenating the patient representation without modification. C_LI
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 94%
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 94%
Similar papers in this journal
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 95%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 95%
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 95%
Similar papers in this journal
- scaLR: a low-resource deep neural network-based platform for single cell analysis and biomarker discovery 95%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 95%
- Explainable deep neural networks for predicting sample phenotypes from single-cell transcriptomics 94%
Similar papers in this journal
Similar papers in this journal
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 95%
- MFmap: A semi-supervised generative model matching cell lines to tumours and cancer subtypes 95%
- Two-step multi-omics modelling of drug sensitivity in cancer cell lines to identify driving mechanisms 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.