Back

Machine Learning Driven Simulations of SARS-CoV-2 Fitness Landscape

Durumeric, A. E. P.; Koehler, J.; Elez, K.; Raich, L.; Suriana, P. A.; Sztain, T.

2024-09-23 bioinformatics
10.1101/2024.09.20.614179 bioRxiv
Show abstract

Predicting protein variant effects is a key challenge in preparing for pathogenic viral strains, understanding mutation-linked diseases, and designing new proteins. Protein sequence-structure-function relationships are difficult to model due to complex allosteric and epistatic effects. To investigate efficient modeling strategies, we trained supervised machine learning (ML) models with deep mutational scanning (DMS) libraries of SARS-CoV-2 receptor binding domain (RBD) sequences labeled with angiotensin converting enzyme 2 (ACE2) binding affinity. These models demonstrate superior performance predicting combinatorial mutation effects compared to adding or averaging the effects of point mutations and exhibit strong extrapolative performance ranking omicron variants when training only on wild type (WT) variants. We characterize the RBD fitness landscape combining ML with Markov Chain Monte Carlo simulations to predict evolutionary patterns from the WT sequence, and generate comparable sequence profiles to high fitness sequences in DMS data predicting mutations in unseen omicron variants. These models provide insight into the relationship between RBD sequence elements, and offer a new perspective on the use of DMS to predict emerging viral strains, which we anticipate will be applicable to other evolutionary prediction tasks. To facilitate application and future development of this strategy, we introduce Mavenets: https://github.com/SztainLab/mavenets.

Published in Journal of Chemical Information and Modeling (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.