A method of identifying false positives in the strain-specific variant calling of rice
Kim, S.; Chu, S.-H.; Park, Y.-J.; Lee, C.-Y.
Show abstract
In this study, we investigated the strain-specific effect in genetic variant calling from next-generation sequencing data. For this purpose, we used two major strains of the rice genome, Indica and Japonica, to build different variant calling models that differ in the composition of samples from the two strains. We found that the more the samples differed in their strains from the reference sequence, the more variants were predicted. In particular, the increase in predicted variants was noticeable when the samples that differed in their strains from the reference were included. We used machine learning approaches to understand this finding and compared the performance of different variant calling models using confusion matrices constructed from the predicted variants. We found that a significant proportion of the incrementally predicted variants are potential false positives, which becomes more pronounced the more phylogenetically different accessions from the reference are included in the samples. For the accuracy of the predicted variants, we proposed a method to identify the false positives that can be excluded from the potential false positives if necessary. The proposed method involves calling true variants from the purebred samples. We demonstrated the validity of the proposed method on the different variant calling models and showed a reduction of false positives in the predicted variants. As an example of practical utility, we applied the method to the dbSNP, a database of known variants, and demonstrated a way to identify false positives in the dbSNP. In these respects, this study provides general recommendations for effective practices in strain-specific variant calling in rice.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NGSpop: A desktop software that supports population studies by identifying sequence variations from next-generation sequencing data 95%
- An updated resource of 180K soybean SNP genotyping array based on the T2T reference genome 94%
- Variant calling and genotyping accuracy of ddRAD-seq: comparison with 20X WGS in layers 94%
Similar papers in this journal
Similar papers in this journal
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 94%
- Automatic identification and annotation of MYB gene family members in plants 94%
- High-fidelity (repeat) consensus sequences from short reads using combined read clustering and assembly 94%
Similar papers in this journal
- Variant calling tool evaluation for variable size indel calling from next generation whole genome and targeted sequencing data 95%
- NAVIP: Unraveling the Influence of Neighboring Small Sequence Variants on Functional Impact Prediction 94%
- VarSCAT: A computational tool for sequence context annotations of genomic variants 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.