Back

Detection of fusion transcripts and their genomic breakpoints from RNA sequencing data

Hoogstrate, Y.; Komor, M. A.; Böttcher, R.; van Riet, J.; van de Werken, H. J. G.; van Lieshout, S.; Hoffmann, R.; van den Broek, E.; Bolijn, A. S.; Dits, N.; Sie, D.; van der Meer, D.; Pepers, F.; Bangma, C. H.; van Leenders, A. J. H. L.; Smid, M.; French, P. J.; Martens, J. W. M.; van Workum, W.; van der Spek, P. J.; Janssen, B.; Caldenhoven, E.; Rausch, C.; de Jong, M.; Stubbs, A. P.; Meijer, G. A.; Fijneman, R. J. A.; Jenster, G.

2021-05-17 bioinformatics
10.1101/2021.05.17.441778 bioRxiv
Show abstract

Spliced fusion-transcripts are typically identified by RNA-seq without elucidating the causal genomic breakpoints. However, non poly(A)-enriched RNA-seq contains large proportions of intronic reads spanning also genomic breakpoints. Using 1.274 RNA-seq samples, we investigated what additional information is embedded in non poly(A)-enriched RNA-seq data. Here, we present our novel, graph-based, Dr. Disco algorithm that makes use of both intronic and exonic RNA-seq reads to identify not only fusion transcripts but also genomic breakpoints in gene but also in intergenic regions. Dr. Disco identified TMPRSS2-ERG fusions with genomic breakpoints and other transcribed rearrangements from multiple RNA-sequencing cohorts. In breast cancer and glioma samples Dr. Disco identified rearrangement hotspots near CCND1 and MDM2 and could directly associate this with increased expression. A comparison with matched DNA-sequencing revealed that most genomic breakpoints are not, or minimally, transcribed while also revealing highly expressed translocations missed by DNA-seq. By using the full potential of non poly(A)-enriched RNA-seq data, Dr. Disco can reliably identify expressed genomic breakpoints and their transcriptional effects.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.