Back

Inference with Viral Quasispecies. Methods for Individual Samples Comparative Analysis.

Gregori, J.

2025-02-25 bioinformatics
10.1101/2024.12.30.630765 bioRxiv
Show abstract

The study of viral quasispecies structure and diversity poses distinct challenges when comparing samples, particularly in scenarios involving single observations from individual biosamples collected at different time points or under varying conditions. In such cases, each biosample is summarized by a quasispecies haplotype rank-abundance vector, rendering traditional statistical methods inapplicable due to the lack of replicates. Quasispecies structure indicators are functions of multinomial parameters, whose asymptotic normality is guaranteed by the continuous differentiability of these indicator functions and sufficiently large sample sizes. Accurate variance estimation for these indicators is critical for reliable inference. This study contrasts three approaches for variance estimation: analytical first-order approximations via the Delta method, and variance estimates obtained through Bootstrap resampling (with replacement, infinite population), and through resampling without replacement using the random delete-d jackknife (finite sample). Exact binomial variances of selected proportions are used as benchmarks to calibrate the resampling methods. We analyze the intrinsic sampling properties of each quasispecies indicator using high-depth next-generation sequencing data from a hepatitis C virus (HCV) cell culture experiment. Our results show that while a subset of indicators yield comparable and close variance estimates across methods, metrics highly sensitive to fluctuations in rare haplotype fractions exhibit substantial divergence. Limitations inherent to the resampling methods are discussed. The analytic approach and tests based on normality are finally recommended.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.