Back

Human protein interactome structure prediction at scale with Boltz-2

Ille, A. M.; Markosian, C.; Burley, S. K.; Pasqualini, R.; Arap, W.

2025-08-21 bioinformatics
10.1101/2025.07.03.663068 bioRxiv
Show abstract

In humans, protein-protein interactions mediate numerous biological processes and are central to both normal physiology and disease. While extensive research efforts have aimed to characterize the human protein interactome, atom-scale structural coverage is limited and remains challenging to resolve through experimental methodology alone. Boltz-2, a recent artificial intelligence/machine learning (AI/ML)-based model capable of interaction structure prediction, may serve this experimentally constrained objective. Here, we present de novo computed models of binary human protein interaction structures predicted using Boltz-2 based on biochemically determined interaction data sourced from the IntAct database. We assessed the predicted interaction structures through different confidence metrics, examined annotated protein domains with putative interaction involvement, and uncovered interaction networks within the context of biological processes and cancer, highlighting extensive interaction involvement of E3 ubiquitin-protein ligase Mdm2 and p53, among other proteins. This work demonstrates the utility of Boltz-2 for structural modeling of the human protein interactome while also providing novel functional and disease contextualization, holding broad significance for biomedical research at large. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC="FIGDIR/small/663068v3_ufig1.gif" ALT="Figure 1"> View larger version (72K): org.highwire.dtl.DTLVardef@5666d8org.highwire.dtl.DTLVardef@7a1d3dorg.highwire.dtl.DTLVardef@115c040org.highwire.dtl.DTLVardef@100bbbe_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.