Back

Crowd-sourced benchmarking of single-sample tumour subclonal reconstruction

Salcedo, A.; Tarabichi, M.; Buchanan, A.; Espiritu, S. M. G.; Zhang, H.; Zhu, K.; Yang, T.-H. O.; Leshchiner, I.; Anastassiou, D.; Guan, Y.; Jang, G. H.; Haase, K.; Deshwar, A. G.; Zou, W.; Umar, I.; Dentro, S.; Wintersinger, J. A.; Chiotti, K.; Demeulemeester, J.; Jolly, C.; Scyza, L.; Ko, M.; PCAWG-11 Working Group, ; SMC-Het Participants, ; Wedge, D. C.; Morris, Q. D.; Ellrot, K.; Van Loo, P.; Boutros, P. C.

2022-06-15 bioinformatics
10.1101/2022.06.14.495937 bioRxiv
Show abstract

Tumours are dynamically evolving populations of cells. Subclonal reconstruction algorithms use bulk DNA sequencing data to quantify parameters of tumour evolution, allowing assessment of how cancers initiate, progress and respond to selective pressures. A plethora of subclonal reconstruction algorithms have been created, but their relative performance across the varying biological and technical features of real-world cancer genomic data is unclear. We therefore launched the ICGC-TCGA DREAM Somatic Mutation Calling -- Tumour Heterogeneity and Evolution Challenge. This seven-year community effort used cloud-computing to benchmark 31 containerized subclonal reconstruction algorithms on 51 simulated tumours. Each algorithm was scored for accuracy on seven independent tasks, leading to 12,061 total runs. Algorithm choice influenced performance significantly more than tumour features, but purity-adjusted read-depth, copy number state and read mappability were associated with performance of most algorithms on most tasks. No single algorithm was a top performer for all seven tasks and existing ensemble strategies were surprisingly unable to outperform the best individual methods, highlighting a key research need. All containerized methods, evaluation code and datasets are available to support further assessment of the determinants of subclonal reconstruction accuracy and development of improved methods to understand tumour evolution.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.