Back

Visual-like Template Diffusion: Boosting Single-Sequence Protein Structure Prediction by Adapting Image Diffusion Models

Wang, X.; Zhang, T.; Cui, Z.; Guo, X.; Wang, F.; Wang, Y.; Cai, X.; Zheng, W.

2025-07-05 bioinformatics
10.1101/2025.07.03.662909 bioRxiv
Show abstract

Single-sequence protein structure prediction has drawn increasing attention due to the high computational costs associated with obtaining homologous information. Here, we propose a visual-like 2D geometric template* diffusion method, named TDFold, to generate high-quality pairwise geometries (including pairwise distances and orientations) for achieving accurate and highly efficient single-sequence 3D structure prediction for proteins. Given a protein sequence, TDFold initially generates high-quality inter-residue geometries from a probabilistic diffusion perspective. Since inter-residue geometries can be encoded as multi-channel feature matrices, analogous to image feature maps, we construct an image-level 2D geometric template diffusion module by adapting the stable diffusion (SD) model from text-vision generation to sequencegeometry diffusion for proteins. Subsequently, a lightweight sequencegeometry collaborative learning (SCL) network is constructed to facilitate accurate and efficient protein structure prediction. As a result, TDFold possesses three highlights: (i) better single-sequence prediction performance: TDFold greatly outperforms existing protein language models (PLMs, e.g. ESMFold and OmegaFold) and homology-based methods (e.g. AlphaFold2, AlphaFold3 and RoseTTAFold) on homologyinsufficient datasets such as Orphan and Orphan25, while also achieving promising results on the popular CASP14, CASP15 and CASP16 benchmarks; (ii) low resource consumption: By utilizing the lightweight SCL architecture, the GPU memory consumption of TDFold is generally lower than that of popular methods such as AlphaFold2 and ESMFold; (iii) higher efficiency in training and inference: TDFold can be trained within a week using a single NVIDIA 4090 GPU. Furthermore, the inference time of TDFold is significantly shorter (about 10x to 100x) than that of existing methods (ESMFold, AlphaFold2 and AlphaFold3) for long protein sequences. This work demonstrates the effectiveness of leveraging powerful vision diffusion models to enhance protein 2D geometric template generation, thereby establishing a new paradigm for single-sequence protein structure prediction. It also accelerates protein-related research, particularly for resource-limited universities and academic institutions. The code has been released to speed up biological research.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.