Back

Extended t-cores for the de novo identification of transposable elements and other inexact repeats from short read RNAseq data

Darmon, S.; Mary, A.; Lacroix, V.

2026-07-10 bioinformatics
10.64898/2026.07.06.736737 bioRxiv
Show abstract

Transcribed repeats represent a major challenge in the de novo assembly of transcriptomes from short RNA-seq reads. Young transposable elements (TEs) and other inexact repeats create dense and ambiguous regions in the assembly graph, preventing the correct assembly of transcripts. In this paper, we introduce a fully de novo method based on the discovery of dense regions in the compacted De Bruijn graph (DBG) to identify such repeats directly from short reads RNA-seq data, without requiring a reference genome or repeat database. Our approach defines the extended t-cores, subgraphs of the DBG that capture the complex topology induced by highly expressed inexact repeats appearing in RNA-seq reads. Independently of its interest for transcriptome assembly, the proposed method appears to be effective for the de novo identification of repeats in transcriptomes. After classifying cores using sequence-based motifs to distinguish simple repeats from potential TEs, we demonstrate its potential for the de novo discovery of transposable elements. We validate the approach on a Mus musculus dataset using expressed TE consensus sequences, showing that extended t-cores correspond to known expressed TE families. We also illustrate its de novo discovery potential on a non-model species, Canis lupus familiaris, where the method was also able to recover known transposable elements.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
14.9%
2
Journal of Computational Biology
48 papers in training set
Top 0.1%
10.9%
3
BMC Genomics
406 papers in training set
Top 0.3%
10.5%
4
PLOS Computational Biology
1863 papers in training set
Top 4%
9.5%
5
Nucleic Acids Research
1281 papers in training set
Top 3%
7.2%
50% of probability mass above
6
BMC Bioinformatics
457 papers in training set
Top 2%
5.4%
7
Genome Biology
637 papers in training set
Top 3%
4.4%
8
NAR Genomics and Bioinformatics
242 papers in training set
Top 1.0%
4.3%
9
Mobile DNA
31 papers in training set
Top 0.1%
3.4%
10
Nature Communications
5641 papers in training set
Top 36%
3.2%
11
Scientific Reports
3612 papers in training set
Top 45%
2.4%
12
Bioinformatics Advances
203 papers in training set
Top 2%
2.4%
13
Genome Research
468 papers in training set
Top 3%
1.9%
14
Genes
144 papers in training set
Top 2%
1.7%
15
Frontiers in Genetics
230 papers in training set
Top 3%
1.5%
16
GigaScience
212 papers in training set
Top 3%
1.5%
17
PeerJ
308 papers in training set
Top 7%
1.3%
18
Peer Community Journal
281 papers in training set
Top 5%
1.0%
19
iScience
1154 papers in training set
Top 32%
0.9%
20
PLOS ONE
5266 papers in training set
Top 62%
0.8%
21
Plant Physiology
238 papers in training set
Top 3%
0.6%
22
Genome Biology and Evolution
338 papers in training set
Top 4%
0.6%
23
Algorithms for Molecular Biology
17 papers in training set
Top 0.2%
0.6%