Back

Physicochemical, functional, and evolutionary characteristics of protein loop regions in human and Escherichia coli proteomes

Zhang, L.; Nishi, H.

2023-06-11 bioinformatics
10.1101/2023.06.09.544332 bioRxiv
Show abstract

Protein loops often play crucial roles in the formation of binding and enzyme active sites. However, the general structural and biological characteristics of these loops remain unclear. In this study, we investigated protein loop regions on a large scale from structural and evolutionary perspectives. After removing redundancy at the protein chain level, 555,516, 102,901, and 24,818 loops were extracted from the entire PDB, Homo sapiens, and Escherichia coli proteins, respectively. Regardless of whether they were isolated from humans or E. coli, numerous loop sequences tended to be unique among proteins or protein chains. However, loop properties exhibited high similarity or conservation, including length, distance, and stretch. The CATH classification analysis suggested that most loops connected the same superfamily, while the heterogeneity of superfamily context repertoires was not explained at the topology or homologous level. In contrast, the functions of conserved loops between human and E. coli proteins were not consistently conserved, with sequences exhibiting considerable divergence in enrichment. The amino acid composition profiles showed that loops from humans exhibited a preference for serine, whereas E. coli loops had biases for glycine and alanine. Although the amino acid composition was primarily determined by the species, the composition of certain special types of loops clustered separately from other classes, suggesting the existence of conserved loops with complex functions. Collectively, this study provides a detailed overview of protein loops from the structural, functional, and evolutionary perspectives and a vast natural loop repertoire for mining additional information.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.