Inference with Viral Quasispecies. Methods for Individual Samples Comparative Analysis.
Gregori, J.
Show abstract
The study of viral quasispecies structure and diversity poses distinct challenges when comparing samples, particularly in scenarios involving single observations from individual biosamples collected at different time points or under varying conditions. In such cases, each biosample is summarized by a quasispecies haplotype rank-abundance vector, rendering traditional statistical methods inapplicable due to the lack of replicates. Quasispecies structure indicators are functions of multinomial parameters, whose asymptotic normality is guaranteed by the continuous differentiability of these indicator functions and sufficiently large sample sizes. Accurate variance estimation for these indicators is critical for reliable inference. This study contrasts three approaches for variance estimation: analytical first-order approximations via the Delta method, and variance estimates obtained through Bootstrap resampling (with replacement, infinite population), and through resampling without replacement using the random delete-d jackknife (finite sample). Exact binomial variances of selected proportions are used as benchmarks to calibrate the resampling methods. We analyze the intrinsic sampling properties of each quasispecies indicator using high-depth next-generation sequencing data from a hepatitis C virus (HCV) cell culture experiment. Our results show that while a subset of indicators yield comparable and close variance estimates across methods, metrics highly sensitive to fluctuations in rare haplotype fractions exhibit substantial divergence. Limitations inherent to the resampling methods are discussed. The analytic approach and tests based on normality are finally recommended.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Primary case inference in viral outbreaks through analysis of intra-host variant population 94%
- Profile hidden Markov model sequence analysis can help remove putative pseudogenes from DNA barcoding and metabarcoding datasets 93%
- Comprehensive benchmarking of metagenomic classification tools for long-read sequencing data 93%
Similar papers in this journal
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 94%
- Short k-mer Abundance Profiles Yield Robust Machine Learning Features and Accurate Classifiers for RNA Viruses 93%
- Estimating multiplicity of infection, haplotype frequencies, and linkage disequilibria from multi-allelic markers for molecular disease surveillance 93%
Similar papers in this journal
- MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores 93%
- Clonal reconstruction from time course genomic sequencing data 92%
- Machine learning based imputation techniques for estimating phylogenetic trees from incomplete distance matrices 92%
Similar papers in this journal
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 94%
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 93%
- Over-optimism in unsupervised microbiome analysis: Insights from network learning and clustering 93%
Similar papers in this journal
- Comparing full variation profile analysis with the conventional consensus method in SARS-CoV-2 phylogeny 94%
- Comprehensive evaluation of methods for differential expression analysis of metatranscriptomics data 93%
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.