Template-based RNA structure prediction advanced through a blind code competition
Lee, Y.; He, S.; Oda, T.; Rao, G. J.; Kim, Y.; Kim, R.; Kim, H.; Heng, C. K.; Kowerko, D.; Li, H.; Nguyen, H.; Sampathkumar, A.; Enrique Gomez, R.; Chen, M.; Yoshizawa, A.; Kuraishi, S.; Ogawa, K.; Zou, S.; Paullier, A.; Zhao, B.; Chen, H.-L.; Hsu, T.-A.; Hirano, T.; Gezelle, J. G.; Haack, D.; Hong, Y.; Jadhav, S.; Koirala, D.; Kretsch, R. C.; Lewicka, A.; Li, S.; Marcia, M.; Piccirilli, J.; Rudolfs, B.; Srivastava, Y.; Steckelberg, A.-L.; Su, Z.; Toor, N.; Wang, L.; Yang, Z.; Zhang, K.; Zou, J.; Baker, D.; Chen, S.-J.; Chiu, W.; Demkin, M.; Favor, A.; Hummer, A. M.; Joshi, C. K.; Kryshtafovyc
Show abstract
Automatically predicting RNA 3D structure from sequence remains an unsolved challenge in biology and biotechnology. Here, we describe a Kaggle code competition engaging over 1700 teams and 43 previously unreleased structures to tackle this challenge. The top three submitted algorithms achieved scores within statistical error of the winners of the recent CASP16 competition. Unexpectedly, the top Kaggle strategy involved a pipeline for discovering 3D templates, without the use of deep learning. We integrated this template-modeling pipeline and other Kaggle strategies to develop a single model RNAPro that retrospectively outperformed individual Kaggle models on the same test set. These results suggest a growing importance of template-based modeling in RNA structure prediction.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- US-align: Universal Structure Alignments of Proteins, Nucleic Acids, and Macromolecular Complexes 95%
- CryoSTAR: Leveraging Structural Prior and Constraints for Cryo-EM Heterogeneous Reconstruction 95%
- Absolute quantitative and base-resolution sequencing reveals comprehensive landscape of pseudouridine across the human transcriptome 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.