Back

Large-scale structure prediction of DUF-containing protein-protein interactions

Riepenhausen, L.; Costa, F.; Andreeva, A.; Bateman, A.

2026-08-20 bioinformatics
10.64898/2026.08.19.745780 bioRxiv
Show abstract

Motivation: Continuing advances in genome and metagenome sequencing expand the number of identified conserved protein families that remain functionally uncharacterized and contain domains of unknown function (DUFs). Functional-association resources such as STRING provide biological context, but mostly do not distinguish indirect association from physical interaction. We assessed whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins. Results: We generated four structural-prediction cohorts from STRING associations involving DUF-containing proteins and evaluated the predicted complexes using interface ipSAE, average pLDDT and buried surface area. An L2-regularized logistic regression model was trained on an initial cohort of predictions from high-confidence STRING associations to prioritize DUF-containing candidates likely to produce structurally confident AlphaFold 3 complexes. The model was then applied across all 12,535 organisms represented in STRING v12.0, followed by grouping into DUF-family and partner-architecture modules, covering 2,076 unique DUF families. The final L2-model screen contained 12,298 successfully modelled protein pairs, including 1,208 (9.82%) complexes meeting a strict-confidence criterion and 2,433 (19.78%) meeting a more liberal confidence criterion. Two examples suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation. Availability and implementation: Predicted structures and associated metadata are available through Zenodo at https://doi.org/10.5281/zenodo.21875362. The model implementation and code used to generate the analyses and figures are available at https://github.com/linoriep/Proteome-scale-structure-prediction-of-DUF-containing-protein-protein-interactions.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.