Strand Orientation Bias Detector (SOBDetector) to remove FFPE sequencing artifacts
Diossy, M.; Sztupinszki, Z.; Krzystanek, M.; Eklund, A. C.; Csabai, I.; Pedersen, A. G.; Szallasi, Z.
Show abstract
Formalin-fixed paraffin-embedded (FFPE) tissue, the most common tissue specimen stored in clinical practice, presents challenges in the analysis due to formalin-induced artifacts. Here we present SOBDetector, that is designed to remove these artifacts from a list of variants, which were extracted from next-generation sequencing data, using the posterior predictive distribution of a Bayesian logistic regression model trained on TCGA whole exomes. Based on a concordance analysis, SOBDetector outperforms the currently publicly available FFPE-filtration techniques (Accuracy: 0.9 {+/-} 0.015, AUC = 0.96). SOBDetector is implemented in Java 1.8 and is freely available online.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RNAIndel: a machine-learning framework for discovery of somatic coding indels using tumor RNA-Seq data 95%
- diffONT: predicting methylation-specific PCR biomarkers based on nanopore sequencing data for clinical application 94%
- Assigning mutational signatures to individual samples and individual somatic mutations with SigProfilerAssignment 94%
Similar papers in this journal
Similar papers in this journal
- DeNovoCNN: A deep learning approach to de novo variant calling in next generation sequencing data 95%
- OmicsFootPrint: a framework to integrate and interpret multi-omics data using circular images and deep neural networks 94%
- Discovering single nucleotide variants and indels from bulk and single-cell ATAC-seq 94%
Similar papers in this journal
Similar papers in this journal
- SUITOR: selecting the number of mutational signatures through cross-validation 95%
- Efficient and Flexible Integration of Variant Characteristics in Rare Variant Association Studies Using Integrated Nested Laplace Approximation 94%
- Epiclomal: probabilistic clustering of sparse single-cell DNA methylation data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.