From partial to whole genome imputation of SARS-CoV-2 for epidemiological surveillance
Ortuno, F. M.; Loucera, C.; Casimiro-Soriguer, C. S.; Lepe, J. A.; Camacho Martinez, P.; Merino Diaz, L.; Chueca, N.; de Salazar, A.; Garcia, F.; Perez-Florido, J.; Dopazo, J.
Show abstract
Backgroundthe current SARS-CoV-2 pandemic has emphasized the utility of viral whole genome sequencing in the surveillance and control of the pathogen. An unprecedented ongoing global initiative is increasingly producing hundreds of thousands of sequences worldwide. However, the complex circumstances in which viruses are sequenced, along with the demand of urgent results, causes a high rate of incomplete and therefore useless, sequences. However, viral sequences evolve in the context of a complex phylogeny and therefore different positions along the genome are in linkage disequilibrium. Therefore, an imputation method would be able to predict missing positions from the available sequencing data. ResultsWe developed impuSARS, an application that includes Minimac, the most widely used strategy for genomic data imputation and, taking advantage of the enormous amount of SARS-CoV-2 whole genome sequences available, a reference panel containing 239,301 sequences was built. The impuSARS application was tested in a wide range of conditions (continuous fragments, amplicons or sparse individual positions missing) showing great fidelity when reconstructing the original sequences. The impuSARS application is also able to impute whole genomes from commercial kits covering less than 20% of the genome or only from the Spike protein with a precision of 0.96. It also recovers the lineage with a 100% precision for almost all the lineages, even in very poorly covered genomes (< 20%) Conclusionsimputation can improve the pace of SARS-CoV-2 sequencing production by recovering many incomplete or low-quality sequences that would be otherwise discarded. impuSARS can be incorporated in any primary data processing pipeline for SARS-CoV-2 whole genome sequencing.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SARS-CoV-2 surveillance in Italy through phylogenomic inferences based on Hamming distances derived from functional annotations of SNPs, MNPs and InDels 96%
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 94%
- Cov2clusters: genomic clustering of SARS-CoV-2 sequences 93%
Similar papers in this journal
Similar papers in this journal
- Development of a nextflow bioinformatics pipeline for the detection of SARS-CoV-2 co-infection cases from genomic surveillance in the Philippines 94%
- Analysis of the ARTIC V4 and V4.1 SARS-CoV-2 primers and their impact on the detection of Omicron BA.1 and BA.2 lineage defining mutations 94%
- Phylogenomics and population genomics of SARS-CoV-2 in Mexico reveals variants of interest (VOI) and a mutation in the Nucleocapsid protein associated with symptomatic versus asymptomatic carriers 93%
Similar papers in this journal
- Machine learning-based approach KEVOLVE efficiently identifies SARS-CoV-2 variant-specific genomic signatures 95%
- Performance of amplicon and capture based next-generation sequencing approaches for the epidemiological surveillance of Omicron SARS-CoV-2 and other variants of concern. 94%
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 93%
Similar papers in this journal
- Intrahost SARS-CoV-2 k-mer identification method (iSKIM) for rapid detection of mutations of concern reveals emergence of global mutation patterns 96%
- Molecular Genetic Analysis of SARS-CoV-2 Lineages in Armenia 94%
- The evolutionary landscape of SARS-CoV-2 variant B.1.1.519 and its clinical impact in Mexico City 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.