Back

FoldToken3: Fold Structures Worth 256 Words or Less

Gao, Z.; Tan, C.; Li, S. Z.

2024-07-08 molecular biology
10.1101/2024.07.08.602548 bioRxiv
Show abstract

Protein structure tokenization has attracted increasing attention in both protein representation learning and generation. While recent work, like FoldToken2 and ESM3, has achieved good reconstruction performance, the compressoin ratio is still limited. In this work, we propose FoldToken3, a novel protein structure tokenization method that can compress protein structures into 256 tokens or less and ensure the reconstruction quality comparable to FoldToken2. To the best of our knowledge, FoldToken3 is the most efficient, light-weight, and compression-friendly protein structure tokenization method. And it will benifit a wide range of protein structure-related tasks, such as protein structure alignment, generation, and representation learning. The work is still in progress and the code will be available upon acceptance.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.