Benchmarking computational methods for multi-omics biomarker discovery in cancer
Li, A. Z.; Du, Y.; Liu, Y.; Chen, L.; Liu, R.
Show abstract
Multi-omics profiling characterizes cancer biology and supports biomarker discovery for prognosis and therapy selection. Although numerous computational multi-omics biomarker identification methods have been proposed, their ability to identify clinically relevant biomarkers has not been systematically evaluated, leaving it unclear whether the resulting biomarker nominations are reliable for downstream validation. Here we systematically benchmark 20 representative statistical, machine learning and deep learning methods using curated gold-standard prognostic and therapeutic biomarkers across five real-world datasets. We evaluate performance in terms of both biomarker identification accuracy and stability. Overall, DeePathNet and Deep-KEGG achieve the best performance. Across methods, effective biomarker recovery is associated with the integration of biological knowledge, global feature interactions, multivariate feature attribution, and effective regularization. Analysis of omics type contributions reveals method- and modality-specific biases, highlighting the importance of broader omics integration. We further evaluate methods on simulated datasets to probe sensitivity with controlled signal and noise. By aggregating results from top-performing methods, we construct consensus biomarker panels that nominate candidates for potential investigations. Finally, we provide user-friendly interfaces to allow researchers to benchmark new methods against the 20 baselines or apply selected methods for biomarker identification on custom multi-omics datasets. Our benchmark is publicly available at https://github.com/athanzli/CancerMOBI-Bench.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Community assessment of methods to deconvolve cellular composition from bulk gene expression 96%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 95%
- LEOPARD: missing view completion for multi-timepoints omics data via representation disentanglement and temporal knowledge transfer 95%
Similar papers in this journal
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 95%
- Identifying tumor cells at the single cell level 95%
- A benchmark of computational methods for correcting biases of established and unknown origin in CRISPR-Cas9 screening data 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.