Inferring viral occurrence patterns through a synthetic data simulation
Pimenoff, V. N.; Cleries, R.
Show abstract
Viruses infecting humans are manifold and several of them provoke significant morbidity and mortality. Simulations creating large synthetic datasets from observed multiple viral strain infections in a limited population sample can be a powerful tool to infer significant pathogen occurrence and interaction patterns, particularly if limited number of observed data units is available. Here, to demonstrate diverse human papillomavirus (HPV) strain occurrence patterns, we used log-linear models combined with Bayesian framework for graphical independence network (GIN) analysis. That is, to simulate datasets based on modeling the probabilistic associations between observed viral data points, i.e different viral strain infections in a set of population samples. Our GIN analysis outperformed in precision all oversampling methods tested for simulating large synthetic viral strain-level prevalence dataset from observed set of HPVs data. Altogether, we demonstrate that network modeling is a potent tool for creating synthetic viral datasets for comprehensive pathogen occurrence and interaction pattern estimations.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Gini coefficients for measuring the distribution of sexually transmitted infections among individuals with different levels of sexual activity 92%
- covid19.Explorer : A web application and R package to explore United States COVID-19 data 90%
- HiLDA: a statistical approach to investigate differences in mutational signatures 90%
Similar papers in this journal
- Inferring Tumor Progression in Large Datasets 92%
- Estimation of the force of infection and infectious period of skin sores in remote Australian communities using interval-censored data 92%
- ScTree: Scalable and robust mechanistic integration of epidemiological and genomic data for transmission tree inference 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.