Probabilistic Approach to Understand Errors in Sequencing and its based Applications
Singh, A.
Show abstract
Next Generation Sequencing has been applied in many areas of biology, including quantification of gene expression, Genome-Wide Association Study (GWAS), gene finding, Motif discovery and much more. Due to the vast area of application and importance in key findings in the field, massive genomic data is being generated using high throughput sequencing. Therefore, sequencing quality also needs to be evaluated in light of different applications. In order to develop effective diagnostic and therapeutic approaches, we need to accurately characterize and identify sequencing errors and distinguish these errors from their true genetic variant in sequencing, i.e. misreads follow a binomial distribution and it further can be approximated to the Poisson process for longer sequences. However, the insertion and deletion rates are 1000 times lower than substitution error rates and, therefore, less significant. The model assumes that error arrival at a position is not dependent on an error at other positions. Furthermore, errors in sequences can cause an error in studies based on multiple sequences and they also follow Binomial - Poisson Distribution (for example - Alignment is a merging of two Binomial processes for short sequences and it further can be approximated to Poisson for long sequences (for example - genomic sequence). It provides a systematic way to evaluate the accuracy in sequencing-based applications. Many error suppressing algorithms or techniques are there, and our Binomial Poisson model can provide a further systematic understanding of error behavior in short sequences so that more techniques for error removal can be developed with much efficient suppression rates.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 96%
- Comparing protein-protein interaction networks of SARS-CoV-2 and (H1N1) influenza using topological features 95%
- Principal Component Analysis applied directly to Sequence Matrix 94%
Similar papers in this journal
Similar papers in this journal
- Penguin: A Tool for Predicting Pseudouridine Sites in Direct RNA Nanopore Sequencing Data 92%
- HLA-DR4Pred2: an improved method for predicting HLA-DRB1*04:01 binders 91%
- RevGraphVAMP: A protein molecular simulation analysis model combining graph convolutional neural networks and physical constraints 89%
Similar papers in this journal
- SNPector: SNP inspection tool for diagnosing gene pathogenicity and drug response in a naked sequence 92%
- Analyses of spike protein from first deposited sequences of SARS-CoV2 from West Bengal, India 90%
- Glibenclamide, ATP and Metformin Increases the Expression of Human Bile Salt Export Pump ABCB11 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.