SurvBoard: Standardised Benchmarking for Multi-omics Cancer Survival Models
Wissel, D.; Janakarajan, N.; Grover, A.; Toniato, E.; Martinez, M. R.; Boeva, V.
Show abstract
Multi-omics data, which include genomic, transcriptomic, epigenetic, and proteomic data, are gaining increasing importance for determining the clinical outcomes of cancer patients. Several recent studies have evaluated various multi-modal integration strategies for cancer survival prediction, highlighting the need for standardizing model performance results. Addressing this issue, we introduce SurvBoard, a benchmark framework that standardizes key experimental design choices. SurvBoard enables comparisons between single-cancer and pan-cancer data models and assesses the benefits of using patient data with missing modalities. We also address common pitfalls in preprocessing and validating multi-omics cancer survival models. We apply SurvBoard to several exemplary use cases, further confirming that statistical models tend to outperform deep learning methods, especially for metrics measuring survival function calibration. Moreover, most models exhibit better performance when trained in a pan-cancer context and can benefit from leveraging samples for which data of some omics modalities are missing. We provide a web service for model evaluation and to make our benchmark results easily accessible and viewable: https://www.survboard.science/. All code is available on GitHub: https://github.com/BoevaLab/survboard/. All benchmark outputs are available on Zenodo: https://zenodo.org/records/11066227. O_TEXTBOXKey MessagesO_LIWe introduce SurvBoard, a comprehensive benchmarking framework for the standardized evaluation of multi-omics cancer survival models. SurvBoard provides an easily accessible platform for the reproducible comparison of models trained on single-cancer and pan-cancer datasets. The platform addresses issues such as the impact of missing modalities and variability in experimental setups. SurvBoard integrates data from four major cancer programs-TCGA, ICGC, TARGET, and METABRIC-to ensure a comprehensive evaluation across diverse types of cancer and research centers. C_LIO_LISurvBoard results confirm that statistical models generally outperform deep learning models in survival function calibration. We also find that pan-cancer training enhances model performance and that models benefit from incorporating data with missing modalities. C_LIO_LISurvBoard includes a web service that allows researchers to submit models for benchmarking and evaluation. A leaderboard is accessible via https://survboard.science/ to promote transparency and the continuous assessment of models performance. C_LI C_TEXTBOX
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Deep Survival EWAS approach estimating risk profile based on pre-diagnostic DNA methylation: an application to Breast Cancer time to diagnosis 94%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 93%
- Survival analysis of pathway activity as a prognostic determinant in breast cancer 93%
Similar papers in this journal
- SurvBenchmark: comprehensive benchmarking study of survival analysis methods using both omics data and clinical data 95%
- ShinyLearner: A containerized benchmarking tool for machine-learning classification of tabular data 92%
- An Integrative Multi-Omics Random Forest Framework for Robust Biomarker Discovery 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.