Back

EpiRanha: Hunting for Epitope Similarity with a Structure- and Residue-Aware Graph Neural Network

Francissen, T.; Babukhian, M.; Britze, H.; Wilke, Y.; Spreafico, R.; Demharter, S.; Arts, M.

2026-04-23 bioinformatics
10.64898/2026.04.21.719830 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWPrecise epitope recognition underpins the efficacy and safety of therapeutic antibodies, yet existing approaches to epitope similarity scoring rely largely on sequence identity or rigid structural superposition, limiting their ability to robustly assess cross-reactivity and potential off-target interactions. We introduce EpiRanha, a multimodal framework that integrates residue-level ESM-2 sequence embeddings with an E(n)-equivariant graph neural network operating on three-dimensional protein structure. EpiRanha produces per-residue "fingerprints" that jointly encode sequence-level context and spatial organization, and then applies a beam-search strategy to identify and rank multiple high-confidence epitope candidates across protein surfaces that are similar to a given query epitope. We evaluate EpiRanha against TM-align on nanobody-antigen complexes from SAbDab-nano and a set of AlphaFold-predicted proteins. EpiRanha consistently recovers the query epitope on its cognate antigen, including highly discontiguous conformational epitopes that rigid alignment methods such as TM-align often fail to capture, achieving lower structural loss and fewer false negatives through flexible residue-level mapping. Beyond these self-matches, EpiRanha also identifies biologically plausible epitope-level similarities on other proteins. Overall, EpiRanha advances epitope characterization beyond sequence or geometry alone, enabling more robust off-target risk assessment, informing training-set construction for predictive models, and more selective antibody design.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.