Optimizing Protein Tokenization: Reduced Amino Acid Alphabets for Efficient and Accurate Protein Language Models
Rannon, E.; Burstein, D.
Show abstract
Protein language models (pLMs) typically tokenize sequences at the single-amino-acid level using a 20-residue alphabet, resulting in long input sequences and high computational cost. Sub-word tokenization methods such as Byte Pair Encoding (BPE) can reduce sequence length but are limited by the sparsity of long patterns in proteins encoded by the standard amino acid alphabet. Reduced amino acid alphabets, which group residues by physicochemical properties, offer a potential solution but their performances with sub-word tokenization have not been systematically studied. In this work, we investigate the combined use of reduced amino acid alphabets and BPE tokenization in protein language models. We pre-train RoBERTa-based pLMs de novo using multiple reduced alphabets and evaluate them across diverse downstream tasks. Our results show that reduced alphabets enable substantially shorter input sequences and faster training and inference, while maintaining comparable - and in some cases improved - performance relative to models trained on the full 20-amino-acid alphabet. These findings demonstrate that alphabet reduction facilitates more effective sub-word tokenization and provides a favorable trade-off between efficiency and predictive accuracy.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DistilProtBert: A distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts 97%
- Effect of Tokenization on Transformers for Biological Sequences 97%
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 97%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.