Back

Natural language-based representation and modeling of RBP binding

Elhajjajy, S.; Weng, Z.

2026-02-08 bioinformatics
10.64898/2026.02.05.704032 bioRxiv
Show abstract

RNA-binding proteins (RBPs) are critical regulators of the human transcriptome, but the binding patterns of most RBPs are insufficiently characterized. While sequence context facilitates RBP binding specificity, its precise contribution remains unclear. Existing computational methods to decipher RBP binding patterns are limited by their architecture-dependence, challenging interpretability, and, importantly, lack of focus on context. We present a novel comprehensive approach to address the aforementioned knowledge gaps. We first introduce a natural language-based representation to model RNA sequences using lexical, syntactic, and semantic forms, then devise a sequence decomposition method based on these structures to deconstruct RNA sequences into regions, each containing a target k-mer and its flanking contexts. We leverage this linguistic conceptualization to predict RBP binding under a Multiple Instance Learning (MIL) framework, which we solve using a novel method of significant region extraction termed "iterative relabeling". We demonstrate that our bottom-up approach discovers key regions contributing to RBP binding in an architecture-dependent, accurate, and interpretable manner.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.