Investigating Sensitivity, Specificity and Accuracy of Variant Calling Pipelines for Analyzing SARS-CoV-2 Data
Krishna, A.; Choi, J. S.
Show abstract
The rapidly increasing popularity of Next Generation Sequencing and analysis methods in clinical and research settings necessitates an understanding of ideal combinations in identifying genomic variants. Especially with the importance of detecting accurate variants for the development of targeted SARS-CoV-2 vaccines. This research compares the results of two Mapping Algorithms , BWA-MEM and Bowtie2, and two Variant Calling Algorithms , LoFreq and FreeBayes, and their combinatory Variant Calling Pipelines on the analyses of Next Generation Sequencing (NGS) data of five SARS-CoV-2 samples collected from patients in the USA, India, Italy, and Malawi and sourced for this research from the publicly available NCBI SRA database. Our analysis of mapping algorithms found that BWA-MEM likely has higher sensitivity and specificity than Bowtie2 for mapping reads, and their specificity and sensitivity vary with read length. Furthermore, the accuracy of variant calling algorithms increases with the number of reads, while higher read length possibly leads to divergence in accuracy and sensitivity. Overall, FreeBayes was found to likely be more sensitive to detecting variants when used with Bowtie2 rather than BWA-MEM for analyzing SARS-CoV-2 data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- In silico comparative genomics of SARS-CoV-2 to determine the source and diversity of the pathogen in Bangladesh 95%
- Natural field diagnosis and molecular confirmation of fungal and bacterial watermelon pathogens in Bangladesh: A case study from Natore and Sylhet district 95%
- Comparative Genomic Study for Revealing the Complete Scenario of COVID-19 Pandemic in Bangladesh 95%
Similar papers in this journal
- Climate influences scrub typhus occurrence in Vellore, Tamil Nadu, India: Analysis of a 15 year dataset 95%
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 94%
- Principal Component Analysis applied directly to Sequence Matrix 94%
Similar papers in this journal
- Extensive In Silico Analysis of the Functional and Structural Consequences of SNPs in Human ARX Gene 95%
- Epitope-Based Peptide Vaccine against Bombali Ebolavirus Viral Protein 40: An Immunoinformatics Combined with Molecular Docking Studies 95%
- Viral miRNAs Confer Survival in Host Cells by Targeting Apoptosis Related Host Genes 94%
Similar papers in this journal
- An Issue of Concern: Unique Truncated ORF8 Protein Variants of SARS-CoV-2 95%
- A machine learning approach for identification of gastrointestinal predictors for the risk of COVID-19 related hospitalization 94%
- Identification of novel mutations in RNA-dependent RNA polymerases of SARS-CoV-2 and their implications on its protein structure 94%
Similar papers in this journal
- Mutational analysis and assessment of its impact on proteins of SARS-CoV-2 genomes from India 95%
- Improving the phylogenetic resolution of Malaysian and Javan mahseer (Cyprinidae), Tor tambroides and Tor tambra: Whole mitogenomes sequencing, phylogeny and potential mitogenome markers 93%
- Systems Biology under heat stress in Indian Cattle 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.