An unsupervised framework for comparing SARS-CoV-2 protein sequences using LLMs
Littlefield, S. B.; Campbell, R. H.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThe severe acute respiratory system coronavirus 2 (SARS-CoV-2) pandemic led to 700 million infections and 7 million deaths worldwide. While studying these viruses, scientists developed a large amount of sequencing data that was made available to researchers. Large language models (LLMs) are pre-trained on large databases of proteins and prior work has shown its use in studying the structure and function of proteins. This paper proposes an unsupervised framework for characterizing SARS-CoV-2 sequences using large language models. First, we perform a comparison of several protein language models previously proposed by other authors. This step is used to determine how clustering and classification approaches perform on SARS-CoV-2 sequence embeddings. In this paper, we focus on surface glycoprotein sequences, also known as spike proteins in SARS-CoV-2 because scientists have previously studied their involvement in being recognized by the human immune system. Our contrastive learning framework is trained in an unsupervised manner, leveraging the Levenshtein distance from pairwise alignment of sequences when the contrastive loss is computed by the Siamese Neural Network. The final part of this paper focuses on a comparison with a previous approach on a test dataset containing data from the latter part of the pandemic. In the prediction of emerging variants, the proposed LLM-based approach shows an improvement of 0.1914 in terms of the adjusted rand index clustering compared to a previously proposed approach. This shows the potential of applying large-language models to this field.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 97%
- CRISPR-DIPOFF: An Interpretable Deep LearningApproach for CRISPR Cas-9 Off-Target Prediction 96%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 96%
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- A Comparison of Embedding Aggregation Strategies in Drug-Target Interaction Prediction 95%
- Predicting RNA Sequence-Structure Likelihood via Structure-Aware Deep Learning 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.