Bridging genomic gaps: A versatile SARS-CoV-2 benchmark dataset for adaptive laboratory workflows
Zufan, S. E.; Judd, L. M.; Walsh, C. J.; Sait, M. L.; Ballard, S. A.; Kwong, J.; Stinear, T. P.; Seemann, T.; Howden, B. P.
Show abstract
Genomic sequencings adoption in public health laboratories (PHLs) for pathogen surveillance is innovative yet challenging, particularly in the realm of bioinformatics. Low- and middle-income countries (LMICs) face increased difficulties due to supply chain volatility, workforce training, and unreliable infrastructure such as electricity and internet services. These challenges also extend to high-income countries (HICs) where bioinformatics is nascent in PHLs and hampered by a lack of specialized skills and computational infrastructure. This underlines the urgency for flexible and resource-aware strategies in genomic sequencing to improve global pathogen surveillance. In response to these challenges, the present research was conducted to identify and analyse key variables influencing the quality and accuracy of amplicon sequence data. An extensive benchmark dataset was developed that encompassed a diverse collection of isolates, viral loads, primer schemes, library preparation methods, sequencing technologies, and basecalling models, totalling 750 sequences. This dataset was analysed with bioinformatic workflows selected for varying levels of technical capacity. The evaluation focused on quality metrics, consensus accuracy, and common genomic epidemiological indicators. The analysis uncovers complex interactions between multiple parameters in laboratory and bioinformatic processes. emphasising resource-constrained PHLs, practical guidelines are proposed. Insights from the benchmark dataset aim to guide the establishment of specific laboratory and bioinformatics protocols for amplicon sequencing in these settings. The findings can also be used to guide the creation of specialised training curricula, further advancing genomic equity. The benchmark dataset itself allows laboratories to customise and evaluate workflows, catering to their distinct requirements and capacities. Such a holistic approach is imperative to build the capacity to monitor pathogens worldwide. Author summaryThis study marks a step toward equity in the field of pathogen genomics, especially for resource-constrained PHLs. It develops and evaluates a comprehensive amplicon sequencing benchmark dataset, offering vital insights for PHLs engaged in genomic surveillance. In particular, the study finds that the choice of basecaller model has a minimal impact on the quality and accuracy of consensus sequences derived from ONT data, which is crucial for labs with limited computational resources. It also highlights the effectiveness of longer amplicons in ensuring consistent coverage and reducing amplicon dropouts at higher viral loads. While Illumina remains a gold standard for data quality, the combination of the Midnight primer scheme with ONTs Rapid library preparation is shown to be a viable alternative, reducing costs, procedural complexity, and hands-on time. The study synthesises these findings into practical guidelines to aid in the development of amplicon sequencing workflows for SARS-CoV-2 with implications for other pathogens.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Analysis of the ARTIC V4 and V4.1 SARS-CoV-2 primers and their impact on the detection of Omicron BA.1 and BA.2 lineage defining mutations 96%
- Bioinformatic investigation of discordant sequence data for SARS-CoV-2: insights for robust genomic analysis during pandemic surveillance 95%
- Identifying the best PCR enzyme for library amplification in NGS 94%
Similar papers in this journal
- Performance of amplicon and capture based next-generation sequencing approaches for the epidemiological surveillance of Omicron SARS-CoV-2 and other variants of concern. 97%
- Nanopore Sequencing of SARS-CoV-2: Comparison of Short and Long PCR-tiling Amplicon Protocols 96%
- Sequencing SARS-CoV-2 from Antigen Tests 96%
Similar papers in this journal
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 94%
- Flexible, Production-Scale, Human Whole Genome Sequencing On A Benchtop Sequencer 94%
- MicroPIPE: An end-to-end solution for high-quality complete bacterial genome construction 94%
Similar papers in this journal
- Rapid, High-Throughput, Cost Effective Whole Genome Sequencing of SARS-CoV-2 Using a Condensed One Hour Library Preparation of the Illumina DNA Prep Kit 96%
- SARS-CoV-2 Genome Sequencing Methods Differ In Their Ability To Detect Variants From Low Viral Load Samples 96%
- Performance of COVIDSeq and Swift normalase amplicon SARS-CoV-2 panels for SARS-CoV-2 Genomes Sequencing: Practical Guide and Combining FASTQ Strategy 95%
Similar papers in this journal
- Benchmark of thirteen bioinformatic pipelines for metagenomic virus diagnostics using datasets from clinical samples 95%
- Multicenter benchmarking of short and long read wet lab protocols for clinical viral metagenomics 95%
- A simplified, amplicon-based method for whole genome sequencing of human respiratory syncytial viruses 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.