Back

DDI_single: Single-Sequence-Based Protein Domain Assembly

Shengyi, Z.

2026-06-08 bioinformatics
10.64898/2026.06.05.730531 bioRxiv
Show abstract

Domains are the basic units of protein structure and function. Appropriate inter-domain organization is critical to enable cooperative execution of multiple related functions. It is thus a crucial step to determine the full-length structure of multi-domain proteins for the purpose of elucidating their functions and designing new drugs to regulate these functions. Existing structure prediction algorithms are generally better at solving the internal conformation of domains, rather than modeling the relative positions between domains. To address the challenge of accurately determining multi-domain protein conformations, we develop a single-sequence-based domain assembly algorithm called DDI_single. DDI_single directly extracts features from the amino acid sequence using the protein language model ESM-lb, and accurately predicts the interactions between residue pairs of structural domains through a novel gated cross-attention module, thus achieving the correct assembly of structural domains. With the knowledge of domain definition, DDI_single achieves more than 20% higher accuracy in the task of predicting the relative distances of residue pairs between domains than that of the single-sequence-based structure prediction algorithm trRosettaX_single. When assembling domains with known spatial conformations, DDI_single correctly assembles 74.4% of the samples in the test set (TM-score>0.5). When assembling domains with unknown spatial conformations, in cases where the internal spatial conformations of domains are correctly modeled, DDI_single correctly assembles 73.9% of the samples.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
12.8%
2
Nature Communications
5641 papers in training set
Top 18%
9.7%
3
Protein Science
246 papers in training set
Top 0.4%
7.8%
4
Journal of Molecular Biology
232 papers in training set
Top 0.3%
5.5%
5
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.3%
4.3%
6
Cell Systems
201 papers in training set
Top 1%
4.3%
7
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
4.3%
8
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.0%
50% of probability mass above
9
Nature Machine Intelligence
70 papers in training set
Top 0.7%
4.0%
10
Communications Biology
993 papers in training set
Top 7%
2.8%
11
Nature Computational Science
55 papers in training set
Top 0.2%
2.7%
12
PLOS Computational Biology
1863 papers in training set
Top 11%
2.7%
13
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
14
Scientific Reports
3612 papers in training set
Top 54%
1.7%
15
Bioinformatics Advances
203 papers in training set
Top 3%
1.7%
16
Genome Research
468 papers in training set
Top 4%
1.7%
17
Structure
193 papers in training set
Top 2%
1.4%
18
Nature Methods
385 papers in training set
Top 5%
1.3%
19
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 35%
1.1%
20
Cell Discovery
57 papers in training set
Top 0.8%
1.1%
21
PRX Life
42 papers in training set
Top 0.7%
1.1%
22
Journal of Structural Biology
64 papers in training set
Top 0.5%
1.1%
23
Nature Structural & Molecular Biology
18 papers in training set
Top 0.4%
1.1%
24
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
25
Journal of Computational Biology
48 papers in training set
Top 0.9%
1.0%
26
Advanced Science
286 papers in training set
Top 9%
0.9%
27
PLOS ONE
5266 papers in training set
Top 61%
0.8%
28
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.2%
0.8%
29
Molecular Biology and Evolution
542 papers in training set
Top 5%
0.8%
30
iScience
1154 papers in training set
Top 40%
0.6%