(Machine) Learning the mutation signatures of SARS-CoV-2: a primer for predictive prognosis
Nagpal, S.; Pinna, N. K.; Srivastava, D.; Singh, R.; Mande, S. S.
Show abstract
MotivationContinuous emergence of new variants through appearance, accumulation and disappearance of mutations in viruses is a hallmark of many viral diseases. SARS-CoV-2 and its variants have particularly exerted tremendous pressure on global healthcare system owing to their life threatening and debilitating implications. The sheer plurality of the variants and huge scale of genome sequence data available for Covid19 have added to the challenges of traceability of mutations of concern. The latter however provides an opportunity to utilize SARS-CoV-2 genomes and the mutations therein as big data records to comprehensively classify the variants through the (machine) learning of mutation patterns. The unprecedented sequencing effort and tracing of disease outcomes provide an excellent ground for identifying important mutations by developing machine learnt models or severity classifiers using mutation profile of SARS-CoV-2. This is expected to provide a significant impetus to the efforts towards not only identifying the mutations of concern but also exploring the potential of mutation driven predictive prognosis of SARS-CoV-2. ResultsWe describe how a graduated approach of building various severity specific machine learning classifiers, using only the mutation corpus of SARS-CoV-2 genomes, can potentially lead to the identification of important mutations and guide potential prognosis of infection. We demonstrate the applicability of model derived important mutations and use of Shapley values in order to identify the significant mutations of concern as well as for developing sparse models of outcome classification. A total of 77,284 outcome traced SARS-CoV-2 genomes were employed in this study which represented a total corpus of 30346 unique nucleotide mutations and 18647 amino acid mutations. Machine learning models pertaining to graduated classifiers of target outcomes namely Asymptomatic, Mild, Symptomatic/Moderate, Severe and Fatal were built considering the TRIPOD guidelines for predictive prognosis. Shapley values for model linked important mutations were employed to select significant mutations leading to identification of less than 20 outcome driving mutations from each classifier. We additionally describe the significance of adopting a temporal modeling approach to benchmark the predictive prognosis linked with continuously evolving pathogens. A chronologically distinct sampling is important in evaluating the performance of models trained on past data in accurately classifying prognosis linked with genomes of future (observed with new mutations). We conclude that while machine learning approach can play a vital role in identifying relevant mutations, caution should be exercised in using the mutation signatures for predictive prognosis in cases where new mutations have accumulated along with the previously observed mutations of concern. Contactsharmila.mande@tcs.com Supplementary informationSupplementary data are enclosed.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MACI: A machine learning-based approach to identify drug classes of antibiotic resistance genes from metagenomic data 92%
- SARS-CoV-2 NSP14 governs mutational instability and assists in making new SARS-CoV-2 variants 92%
- L-shape distribution of the relative substitution rate (c micro) observed for SARS-COV-2 genome, inconsistent with the selectionist theory, the neutral theory and the nearly neutral theory but a near-neutral balanced selection theory: implication on neutralist-selectionist debate 92%
Similar papers in this journal
- Short k-mer Abundance Profiles Yield Robust Machine Learning Features and Accurate Classifiers for RNA Viruses 94%
- Structural impact of synonymous mutations in six SARS-CoV-2 Variants of Concern 94%
- Machine learning-based approach KEVOLVE efficiently identifies SARS-CoV-2 variant-specific genomic signatures 94%
Similar papers in this journal
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 93%
- Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-Human interactions 93%
- Modeling and analysis of site-specific mutations in cancer identifies known plus putative novel hotspots and bias due to contextual sequences 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.