Back

TEscape: Defining the human transposable element transcriptome using multiplatform long-read sequencing

Mercuri, R. L. V.; Mombach, D. M.; dos Santos, F. R. C.; Perez-Schindler, J.; Huang, Y.; Spealman, P.; Pintacuda, G.; Al'Khafaji, A.; Donnard, E. R.; Claussnitzer, M.; Galante, P. A. F.

2026-07-12 bioinformatics
10.64898/2026.07.08.737305 bioRxiv
Show abstract

Transposable elements (TEs) not only account for half of the human genome sequence but also generate transcripts that contribute to transcriptomic diversity. Yet, their repetitive nature has hindered accurate quantification of the full TE-derived transcriptome, a challenge that long-read sequencing can overcome. Here, we combined multiplexed arrays isoform sequencing (MAS-ISO-seq) with a dedicated computational framework (TEscape) to perform an in-depth annotation of the human TE transcriptome. To capture the breadth of human transcriptome diversity, we profiled six representative cell types spanning three distinct biological contexts, including metabolism with, primary patient-derived adipogenic cells at two differentiation stages, and iPSC derived hepatic progenitor cells; the nervous system with iPSC-derived neurons, neural progenitor cells (NPCs), and pluripotency using induced pluripotent stem cells (iPSCs). Together, these datasets yielded over 235 million full-length long reads. First, to assess data coverage and transcriptome depth, we quantified protein-coding gene expression, detecting 14,312 genes (73.6% of all annotated protein-coding genes), which is a level consistent with deep and comprehensive transcriptome representation. Second, focusing on TE-derived transcripts, we identified >83,000 previously unannotated isoforms, the vast majority (84%) originating from a complex combination of multi-TEs. We also identified solo TEs, which are predominantly from LINE1 (14%). We confirmed that TE-transcripts are able to be exemplified by signatures detected in Liver Hepatocellular Carcinoma (LICH). Together, MAS-ISO-seq and TEscape establish the first long-read-based, high-resolution atlas of transcribed human TEs, providing a foundational resource for integrative transcriptome analyses and for investigating TE expression and regulation in health and disease. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=112 SRC="FIGDIR/small/737305v1_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@18158aeorg.highwire.dtl.DTLVardef@e51fdforg.highwire.dtl.DTLVardef@8f9504org.highwire.dtl.DTLVardef@804113_HPS_FORMAT_FIGEXP M_FIG Graphical Abstract C_FIG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.