Using Genome Sequence Data to Predict SARS-CoV-2 Detection Cycle Threshold Values
Duesterwald, L.; Nguyen, M.; Christensen, P.; Long, S. W.; Olsen, R. J.; Musser, J. M.; Davis, J. J.
Show abstract
The continuing emergence of SARS-CoV-2 variants of concern (VOCs) presents a serious public health threat, exacerbating the effects of the COVID19 pandemic. Although millions of genomes have been deposited in public archives since the start of the pandemic, predicting SARS-CoV-2 clinical characteristics from the genome sequence remains challenging. In this study, we used a collection of over 29,000 high quality SARS-CoV-2 genomes to build machine learning models for predicting clinical detection cycle threshold (Ct) values, which correspond with viral load. After evaluating several machine learning methods and parameters, our best model was a random forest regressor that used 10-mer oligonucleotides as features and achieved an R2 score of 0.521 {+/-} 0.010 (95% confidence interval over 5 folds) and an RMSE of 5.7 {+/-} 0.034, demonstrating the ability of the models to detect the presence of a signal in the genomic data. In an attempt to predict Ct values for newly emerging variants, we predicted Ct values for Omicron variants using models trained on previous variants. We found that approximately 5% of the data in the model needed to be from the new variant in order to learn its Ct values. Finally, to understand how the model is working, we evaluated the top features and found that the model is using a multitude of k-mers from across the genome to make the predictions. However, when we looked at the top k-mers that occurred most frequently across the set of genomes, we observed a clustering of k-mers that span spike protein regions corresponding with key variations that are hallmarks of the VOCs including G339, K417, L452, N501, and P681, indicating that these sites are informative in the model and may impact the Ct values that are observed in clinical samples.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Exact mapping of Illumina blind spots in the Mycobacterium tuberculosis genome reveals platform-wide and workflow-specific biases 95%
- Exploring SNP Filtering Strategies: The Influence of Strict vs Soft Core 94%
- Benchmarking taxonomic classifiers with Illumina and Nanopore sequence data for clinical metagenomic diagnostic applications 94%
Similar papers in this journal
- Targeted Hybridization Capture of SARS-CoV-2 and Metagenomics Enables Genetic Variant Discovery and Nasal Microbiome Insights 96%
- Development of an amplicon nanopore sequencing strategy for detection of mutations conferring intermediate resistance to vancomycin in Staphylococcus aureus strains 93%
- Digital PCR discriminates between SARS-CoV-2 Omicron variants and immune escape mutations 93%
Similar papers in this journal
- A comparison of short- and long-read whole genome sequencing for microbial pathogen epidemiology 94%
- AMiGA: software for automated Analysis of Microbial Growth Assays 93%
- A method to correct for local alterations in DNA copy number that bias functional genomics assays applied to antibiotic-treated bacteria 93%
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 93%
- Nanopore Sequencing of SARS-CoV-2: Comparison of Short and Long PCR-tiling Amplicon Protocols 93%
- A rapid, low cost, and highly sensitive SARS-CoV-2 diagnostic based on whole genome sequencing 93%
Similar papers in this journal
- Predicting Antimicrobial Resistance Using Conserved Genes 95%
- Virus-Host Interactions Predictor (VHIP): machine learning approach to resolve microbial virus-host interaction networks 94%
- Decoding the Language of Microbiomes: Leveraging Patterns in 16S Public Data using Word-Embedding Techniques and Applications in Inflammatory Bowel Disease 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.