Back

Deep-learning predictions of biomolecular structures : persistent limitations and new horizons extended by explicit ion addition

Marien, J.; Sritharan, S.; Caviglia, B.; Versini, R.; Auclair, L.; Barraud, P.; Tisne, C.; Reguei, A.; Murail, S.; Leclerc, F.; Duboue-Dijon, E.; Tao, J.; Basdevant, N.; Baaden, M.; Prevost, C.; Sacquin-Mora, S.; Taly, A.

2026-07-26 biophysics
10.64898/2026.07.24.740587 bioRxiv
Show abstract

The advent of deep learning-driven tools such as AlphaFold has revolutionized the prediction of biomolecular structures, offering unprecedented accuracy and accessibility for proteins, RNA, and their complexes. While these tools have demonstrated remarkable success in benchmarking competitions and enabled experimentalists to generate models with ease, their widespread use has also highlighted persistent challenges. These include difficulties in assessing model confidence, limitations in predicting transmembrane domains, nucleic acids, conformational diversity, and interactions with ions or ligands, as well as the tendency to misfold intrinsically disordered regions (IDRs). In this perspective, we critically evaluate the strengths and limitations of current AI-based structure prediction tools through illustrative examples with a particular emphasis on the impact of explicit ion modelling. We notably report how the explicit addition of a few potassium cations to the prediction of IDRs or G-quadruplexes can trigger massive conformational switches compared to "dry" predictions. On this basis, we suggest modelling sequences both "dry" and in the presence of explicit potassium cations as a simple, practical way to sample alternative conformations and to expose disordered regions that current predictors tend to over-fold. We discuss the importance of reporting confidence metrics in publications to avoid overinterpretation. Furthermore, we address the unique challenges of RNA structure prediction, where data scarcity and structural complexity limit the performance of both classical and deep learning methods. Our analysis underscores the need for continued methodological advancements, integration of complementary computational tools, and expansion of high-quality experimental datasets. TOC Graphic O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/740587v1_ufig1.gif" ALT="Figure 1"> View larger version (16K): org.highwire.dtl.DTLVardef@fc2f42org.highwire.dtl.DTLVardef@82c1a4org.highwire.dtl.DTLVardef@7733cborg.highwire.dtl.DTLVardef@1e97db2_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.