Pangenome References Improve Biomarker Estimation from Tumor Sequencing Data
Arslan, E.; Turgut, D.; Kalay, O.; Demirkaya-Budak, S.; Budak, G.; Jain, A.
Show abstract
It has recently been shown that patients from non-European ancestries are at a higher risk of inappropriate clinical intervention because of inaccurate biomarker estimation, arising from the reference bias inherent in standard methods for determining the tumor genome from sequencing data. Here we demonstrate that these inaccuracies can be reduced by using a pangenome reference appropriate for the patients population. We constructed a novel secondary analysis workflow where the pangenome reference serves as a scaffold for mapping the sequencing reads, and is also included in the relevant panel of normals needed to discriminate between germline and somatic mutations in tumor-only sequencing. This approach detects known somatic mutations in tumor-only sequencing more accurately than the standard GATK somatic calling workflow, prevalent in diagnostic settings for analysis of sequencing data from tumor-only assays, on a standard benchmark tumor sample, HCC1395 (33% relative increase in F1 score). We also assessed the expected clinical impact of our approach by comparing the Tumor Mutational Burden (TMB) calculated from missense somatic mutations called in tumor/normal samples from 6 patients self-reported as belonging to African, 1 to Asian and 3 to European populations respectively. We find that the TMB values calculated from the tumor-only sequencing data analyzed by our workflow more closely approximate the TMB values calculated from the tumor-normal analysis of the same sample, being 35% higher on average, whereas GATK tumor-only analysis generates TMB values 56% higher on average than the tumor-normal analysis of the same sample. Tumor-normal TMB values calculated by the two methods do not vary as drastically, GATK generated values being 13% higher on average, indicating that GATK tumor-only analysis leads to significant overestimation of TMB values, which can be largely corrected by using our workflow when tumor-normal sequencing is not available. These results indicate that pangenome based analysis has the potential to become the new standard for unbiased processing of somatic sequencing samples, following on from its increased adoption for germline sequencing analysis.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sigflow: an automated and comprehensive pipeline for cancer genome mutational signature analysis 94%
- unCOVERApp: an interactive graphical application for clinical assessment of sequence coverage at the base-pair level 93%
- Customized de novo mutation detection for any variant calling pipeline: SynthDNM 93%
Similar papers in this journal
- Accuracy and Reproducibility of Somatic Point Mutation Calling in Clinical-Type Targeted Sequencing Data 94%
- Bioinformatics workflows for genomic analysis of tumors from Patient Derived Xenografts (PDX): challenges and guidelines 93%
- Identification of single nucleotide variants using position-specific error estimation in deep sequencing data 92%
Similar papers in this journal
Similar papers in this journal
- pyCancerSig: subclassifying human cancer with comprehensive single nucleotide, structural and microsatellite mutational signature deconstruction from whole genome sequencing 94%
- Performance analysis of conventional and AI-based variant callers using short and long reads 94%
- Probabilistic modeling methods for cell-free DNA methylation based cancer classification 93%
Similar papers in this journal
- RBV: Read balance validator, a tool for prioritising copy number variations in germline conditions 93%
- ConsensuSV-ONT - a modern method for accurate structural variant calling 93%
- CANCERSIGN: a user-friendly and robust tool for identification and classification of mutational signatures and patterns in cancer genomes 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.