Back

A fully automated approach for quality control of cancer mutations in the era of high-resolution whole genome sequencing

Househam, J.; Cross, W. C.; Caravagna, G.

2021-02-15 bioinformatics
10.1101/2021.02.13.429885 bioRxiv
Show abstract

AbstractThe identification of chromosome number alterations is now widespread in cancer research, but three features of genomic data hinder copy number calling and downstream analyses: the purity of the tumour sample, intra-tumour heterogeneity, and the ploidy of the tumour. To assess these features, consensus methods are often utilised, though these become onerous in projects that involve thousands of genomes. To facilitate the validation of clonal and subclonal copy number variants we present CNAqc, an evolution-inspired toolset that leverages the known quantitative relationships of purity, ploidy and heterogeneity. We validate the algorithms in CNAqc using low-pass single-cell data, as well as extensive simulations. Its application is demonstrated using over 4000 whole genomes and exomes from TCGA, and PCAWG. A real world application of CNAqc in the analysis of clinical tumour samples, has been demonstrated by its incorporation into the validation of clinically accredited bioinformatics pipeline at Genomics England. Our approach is compatible with most bioinformatic pipelines and designed to augment algorithms with automated quality control procedures for data validation.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.