DOMINO: Learning Domain Co-occurrence for Multidomain Protein Design
Dai, F.; Su, J.; Tan, Q.; Yang, H.; Zhou, X.; Yuan, F.
Show abstract
Multidomain proteins arise through the reuse and recombination of structural domains, yet natural architectures represent a sparse, structured sample of the possible domain-combination space. Here, we introduce DOMINO, a two-stage framework that learns domain co-occurrence from TED-annotated multidomain proteins and uses the learned patterns to generate new multidomain sequences. DOMIN, a contrastive retrieval model, embeds domains into a latent compatibility space and retrieves candidate partners for a query domain from a TED-derived domain pool, including pairings not observed in the TED-derived co-occurrence set. DOMO, a conditional autoregressive sequence model, converts each retrieved domain pair into a full-length protein sequence by jointly generating the specified domain regions and the non-domain sequence context between and around them. DOMIN recovers hierarchical patterns of natural domain co-occurrence and expands the observed CATH homologous-superfamily co-occurrence network with candidate novel pairings. DOMO realizes both held-out natural pairs and DOMIN-retrieved pairs as proteins with high domain recovery and high AlphaFold-predicted structural confidence. Applied at scale, DOMINO generated 5 million retrieval-derived multidomain proteins, with sampled designs showing recovery of the specified domains, diverse CATH annotations, and sequence novelty relative to UniRef100. Together, these results support domain co-occurrence as a predictive design prior and demonstrate a scalable strategy for exploring multidomain protein architectures through new combinations of existing structural modules.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DynamicGT: a dynamic-aware geometric transformer model to predict protein binding interfaces in flexible and disordered regions 96%
- Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model 95%
- Undersampling and the inference of coevolution in proteins 95%
Similar papers in this journal
Similar papers in this journal
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 96%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 96%
- Ig-VAE: Generative Modeling of Immunoglobulin Proteins by Direct 3D Coordinate Generation 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.