Back

Comprehensive Database of Circular Permutations: Systematic Detection and Analysis Using Deep Learning

Hu, Y.; Huang, B.

2024-08-28 bioinformatics
10.1101/2024.08.28.610105 bioRxiv
Show abstract

This study developed a robust method to detect circular permutations in the Protein Data Bank, analyzing 287,081 proteins with sequence lengths under 800 residues. By employing Foldseek and MMseqs2 for similarity searches and refining results with TM-align, icarus, and plmCP, we identified 20,801 potential circular permutation pairs and 3,351 unique circular permutation proteins. These findings have been compiled into PermuStructDB, a comprehensive database dedicated to circular permutation proteins. This approach, along with the establishment of PermuStructDB, significantly advances our understanding of protein structural variations and evolutionary adaptations, providing a valuable resource for future research in protein engineering and design.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.