Language modelling for biological sequences - curated datasets and baselines
Almagro Armenteros, J. J.; Johansen, A. R.; Winther, O.; Nielsen, H.
Show abstract
MotivationLanguage modelling (LM) on biological sequences is an emergent topic in the field of bioinformatics. Current research has shown that language modelling of proteins can create context-dependent representations that can be applied to improve performance on different protein prediction tasks. However, little effort has been directed towards analyzing the properties of the datasets used to train language models. Additionally, only the performance of cherry-picked downstream tasks are used to assess the capacity of LMs. ResultsWe analyze the entire UniProt database and investigate the different properties that can bias or hinder the performance of LMs such as homology, domain of origin, quality of the data, and completeness of the sequence. We evaluate n-gram and Recurrent Neural Network (RNN) LMs to assess the impact of these properties on performance. To our knowledge, this is the first protein dataset with an emphasis on language modelling. Our inclusion of properties specific to proteins gives a detailed analysis of how well natural language processing methods work on biological sequences. We find that organism domain and quality of data have an impact on the performance, while the completeness of the proteins has little influence. The RNN based LM can learn to model Bacteria, Eukarya, and Archaea; but struggles with Viruses. By using the LM we can also generate novel proteins that are shown to be similar to real proteins. Availability and implementationhttps://github.com/alrojo/UniLanguage
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 94%
- LMPred: Predicting Antimicrobial Peptides Using Pre-Trained Language Models and Deep Learning 94%
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 95%
- Learning massive interpretable gene regulatory networks of the human brain by merging Bayesian Networks 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.