Accurate Variant Classification in Tumour-Only Genomic Data Using Interpretable Tabular Models
Tattini, L.; Yan, Y.; Chaturvedi, N.; Appuswamy, R.
Show abstract
Recent work has shown that machine learning can provide a reliable tool to classify somatic and rare germline variants in cancer studies where matched-normal samples are not available. Here, we present a workflow that combines an opensource pipeline with three machine-learning models, XGBoost, LightGBM, and TabNet, trained on eight types of features. Our approach substantially enhances the accuracy across all tested models providing accurate results irrespective of sample ancestry and tumour type. We build a parsimonious model and demonstrate that training on low-coverage data retains high accuracy when applied to high-coverage data and vice versa. In contrast to previous findings, our results indicate that XGBoost slightly outperforms LightGBM, achieving high classification accuracy even in the absence of copy-number information and allowing for the ancestry-unbiased calculation of the tumour mutational burden for different types of cancer.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Katdetectr: utilising unsupervised changepoint analysis for robust kataegis detection 96%
- Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design 95%
- cfDNA UniFlow: A unified preprocessing pipeline for cell-free DNA data from liquid biopsies 93%
Similar papers in this journal
Similar papers in this journal
- Efficient and Flexible Integration of Variant Characteristics in Rare Variant Association Studies Using Integrated Nested Laplace Approximation 94%
- Epiclomal: probabilistic clustering of sparse single-cell DNA methylation data 94%
- Revealing cancer driver genes through integrative transcriptomic and epigenomic analyses with Moonlight 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.