Intra-host site-specific polymorphisms of SARS-CoV-2 is consistent across multiple samples and methodologies
Rose, R.; Nolan, D. J.; Moot, S.; Feehan, A.; Cross, S.; Garcia-Diaz, J.; Lamers, S. L.
Show abstract
Despite the potential relevance to clinical outcome, intra-host dynamics of SARS-CoV-2 are unclear. Here, we quantify and characterize intra-host variation in SARS-CoV-2 raw sequence data uploaded to SRA as of 14 April 2020, and compare results between two sequencing methods (amplicon and RNA-Seq). Raw fastq files were quality filtered and trimmed using Trimmomatic, mapped to the WuhanHu1 reference genome using Bowtie2, and variants called with bcftools mpileup. To ensure sufficient coverage, we only included samples with 10X coverage for >90% of the genome (n=406 samples), and only variants with a depth >=10. Derived (i.e. non-reference) alleles were found at 408 sites. The number of polymorphic sites (i.e. sites with multiple alleles) within samples ranged from 0-13, with 72% of samples (295/406) having at least one polymorphic site. Correlation between number of polymorphic sites and coverage was very low for both sequencing methods (R2 < 0.1, p < 0.05). Polymorphisms were observed >1 sample at 66 sites (range: 2-38 samples). The minor allele frequency (MAF) at each shared polymorphic site was 0.03% - 48.5%. 33/66 sites occurred in ORF1a1b, and 37/66 changes were non-synonymous. At 10/66 sites, derived alleles were found in samples sequenced using both methods. Polymorphic amplicon samples were found at 10/10 positions, while polymorphic RNA-Seq samples were found at 7/10 positions. In conclusion, our results suggest that intra-host variation is prevalent among clinical samples. While mutations resulting from amplification and/or sequencing errors cannot be excluded, the observation of shared polymorphic sites with high MAF across multiple samples and sequencing methods is consistent with true underlying variation. Further investigation into intra-host evolutionary dynamics, particularly with longitudinal sampling, is critical for broader understanding of disease progression.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Profiling SARS-CoV-2 mutation fingerprints that range from the viral pangenome to individual infection quasispecies 90%
- A bench-to-data analysis workflow for respiratory syncytial virus whole-genome sequencing with short and long-read approaches 89%
- Evaluating metagenomics and targeted approaches for diagnosis and surveillance of viruses 88%
Similar papers in this journal
- Amplicon and metagenomic analysis of MERS-CoV and the microbiome in patients with severe Middle East respiratory syndrome (MERS) 92%
- Application of a High-Resolution Melt Assay for Monitoring SARS-CoV-2 Variants in Burkina Faso and Kenya 89%
- Conserved Genomic Terminals of SARS-CoV-2 as Co-evolving Functional Elements and Potential Therapeutic Targets 89%
Similar papers in this journal
- Robust clinical detection of SARS-CoV-2 variants by RT-PCR/MALDI-TOF multi-target approach 92%
- Profiling the positive detection rate of SARS-CoV-2 using polymerase chain reaction in different types of clinical specimens: a systematic review and meta-analysis 90%
- Amplification of human β-glucoronidase gene for appraising the accuracy of negative SARS-CoV-2 RT-PCR results in upper respiratory tract specimens 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.