SARITA: A Large Language Model for Generating the S1 Subunit of the SARS-CoV-2 Spike Protein
Rancati, S.; Nicora, G.; Bergomi, L.; Buonocore, T. M.; Czyz, D. M.; Parimbelli, E.; Bellazzi, R.; Salemi, M.; Prosperi, M.; Marini, S.
Show abstract
The COVID-19 pandemic has profoundly impacted global health, economics, and daily life, with over 776 million cases and 7 million deaths from December 2019 to November 2024. Since the original SARS-CoV-2 Wuhan strain emerged, the virus has evolved into variants such as Alpha, Beta, Gamma, Delta, and Omicron, all characterized by mutations in the Spike glycoprotein, critical for viral entry into human cells via its S1 and S2 subunits. The S1 subunit, binding to the ACE2 receptor and mutating frequently, affects infectivity and immune evasion; the more conserved S2, on the other hand, facilitates membrane fusion. Predicting future mutations is crucial for developing vaccines and treatments adaptable to emerging strains, enhancing preparedness and intervention design. Generative Large Language Models (LLMs) are becoming increasingly common in the field of genomics, given their ability to generate realistic synthetic biological sequences, including applications in protein design and engineering. Here we present SARITA, an LLM with up to 1.2 billion parameters, based on GPT-3 architecture, designed to generate high-quality synthetic SARS-CoV-2 Spike S1 sequences. SARITA is trained via continuous learning on the pre-existing protein model RITA. When trained on Alpha, Beta, and Gamma variants (data up to February 2021 included), SARITA correctly predicts the evolution of future S1 mutations, including characterized mutations of Delta, Omicron and Iota variants. Furthermore, we show how SARITA outperforms alternative approaches, including other LLMs, in terms of sequence quality, realism, and similarity with real-world S1 sequences. These results indicate the potential of SARITA to predict future SARS-CoV-2 S1 evolution, potentially aiding in the development of adaptable vaccines and treatments.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 96%
- A deep learning approach to real-time HIV outbreak detection using genetic data 96%
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
- Predicting RNA Sequence-Structure Likelihood via Structure-Aware Deep Learning 94%
Similar papers in this journal
- nanoBERT: A deep learning model for gene agnostic navigation of the nanobody mutational space 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 95%
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.