Back

Widespread DNA off-targeting confounds studies of RNA chromatin occupancy

Goldrich, M. J.; Delhaye, L.; Bekaert, S.-L.; Decaesteker, B.; Van Nieuwerburgh, F.; Speleman, F.; Eyckerman, S.; Mestdagh, P.; Ulitsky, I.

2025-11-11 molecular biology
10.1101/2025.11.11.687850 bioRxiv
Show abstract

The importance of long noncoding RNA (lncRNA) functions is recognized across biological systems, but their modes of action remain poorly understood. One mechanism proposed to be particularly common is gene expression regulation via recruitment to specific genomic regions. Several high-throughput sequencing methods were developed for studying the genome-wide chromatin occupancy of lncRNAs, including ChIRP-seq, CHART-seq, and RAP-seq. These methods utilize biotin-labeled probes targeting the RNA of interest to isolate and recover the chromatin associated with it. Many of the datasets obtained with these methods contain thousands of binding sites, which appears to be in contradiction with the low abundance of the interrogated lncRNAs. We studied the chromatin interactome of NESPR lncRNA in cells with varying levels of endogenous expression and then performed a meta-analysis using dozens of RNA-chromatin interaction datasets in human and mouse cells. We demonstrate that thousands of regions reported to bind lncRNAs most likely arise from the spurious recovery of DNA elements, where the ends of the recovered DNA fragments exhibit partial complementarity with the probes used for the pulldown. In addition, crucial controls were rarely used in studies profiling RNA-chromatin interactions. Therefore, most chromatin regions reported as bound by trans-acting RNAs in recent studies in mammalian cells appear to be technical artifacts. We provide suggestions for assessing the quality of RNA-chromatin datasets and their improvement.

Published in Nature Biotechnology (predicted rank #11) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.