Back

Benchmarking computational methods for multi-omics biomarker discovery in cancer

Li, A. Z.; Du, Y.; Liu, Y.; Chen, L.; Liu, R.

2026-03-14 bioinformatics
10.64898/2025.12.18.695266 bioRxiv
Show abstract

Multi-omics profiling characterizes cancer biology and supports biomarker discovery for prognosis and therapy selection. Although numerous computational multi-omics biomarker identification methods have been proposed, their ability to identify clinically relevant biomarkers has not been systematically evaluated, leaving it unclear whether the resulting biomarker nominations are reliable for downstream validation. Here we systematically benchmark 20 representative statistical, machine learning and deep learning methods using curated gold-standard prognostic and therapeutic biomarkers across five real-world datasets. We evaluate performance in terms of both biomarker identification accuracy and stability. Overall, DeePathNet and Deep-KEGG achieve the best performance. Across methods, effective biomarker recovery is associated with the integration of biological knowledge, global feature interactions, multivariate feature attribution, and effective regularization. Analysis of omics type contributions reveals method- and modality-specific biases, highlighting the importance of broader omics integration. We further evaluate methods on simulated datasets to probe sensitivity with controlled signal and noise. By aggregating results from top-performing methods, we construct consensus biomarker panels that nominate candidates for potential investigations. Finally, we provide user-friendly interfaces to allow researchers to benchmark new methods against the 20 baselines or apply selected methods for biomarker identification on custom multi-omics datasets. Our benchmark is publicly available at https://github.com/athanzli/CancerMOBI-Bench.

Published in Briefings in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.