Systematic comparison of Generative AI-Protein Models reveals fundamental differences between structural and sequence-based approaches.
Barnett, A. J.; KC, R.; Pandey, P.; Somasiri, P.; Fairfax, K. A.; Hewitt, A. W.
10.1101/2025.03.23.644844 bioRxivShow abstract
Recent advances in artificial intelligence have led to the development of generative models for de novo protein design. We compared 13 state-of-the-art generative protein models, assessing their ability to produce feasible, diverse, and novel protein monomers. Structural diffusion models generally create designs with higher confidence in predicted structures and more biologically plausible energy distributions, but exhibit limited diversity and strong sequence biases. Conversely, protein language models generate more diverse and novel designs but with lower structural confidence. We also evaluated these models ability to generate unique proteins, conditionally based on the Tobacco Etch Virus (TEV) protease. Generative models were successful in producing functional enzymes, albeit with diminished activity compared to the wildtype TEV. Our systematic benchmarking provides a foundation for evaluating and selecting generative protein models, while highlighting the complementary strengths of different generative paradigms. This framework will facilitate an informed application of these tools for bio-medical engineering and design.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Towards mechanistic models of mutational effects: Deep Learning on Alzheimer's Aβ peptide 96%
- Systematic Investigation of Machine Learning on Limited Data: A Study on Predicting Protein-Protein Binding Strength 94%
- The Atomic-Level Physiochemical Determinants of T Cell Receptor Dissociation Kinetics 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.