Systematic comparative benchmarking of computational methods for the detection of transposable elements in long-read sequencing data
Seymen, N.; Santos, R.; Lakshmanan, R.; Topp, S.; Al-Chalabi, A.; Al Khleifat, A.; Breen, G.; Dobson, R. J.; Quinn, J. P.; Karimi, M. M.; Iacoangeli, A.
Show abstract
BackgroundMobile element insertions, particularly transposable elements (TEs) such as Alu, LINE-1 (L1), SVA, and endogenous retroviruses (ERVs), represent a major source of human genetic variation and have been implicated in evolution, genomic instability, and disease. Although long-read sequencing generally outperforms short-read sequencing for the characterisation of such elements, their accurate detection with long-reads remains challenging, with different computational tools adopting varying approaches and producing divergent call sets. As gold standards currently do not exist for TE detection, benchmarking these methods is essential to understand their strengths, limitations, and biases. Here, we systematically evaluate the performance of available state-of-the-art TE detection tools on both simulated and real human genome data using highly characterised samples from the Genome in a Bottle consortium, population level reference databases and an in-house collection for which matching short-read sequencing data are available. ResultsOur results show significant differences in calling strategies, leading to substantial variation in precision, recall, and the spectrum of TE families detected across tools. Our benchmark also displays the differences between short-read and long-read calls, highlighting the importance of appropriate method selection. ConclusionsThe benchmarking results presented here will aid TE researchers make better informed decisions on which tool to use in their long-read TE analyses. Strengths and limitations of different tools have been highlighted in depth as well as their computational requirements, which will result in less time spent finding the best tool for the job and promote faster TE research.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comprehensive benchmarking of software for mapping whole genome bisulfite data: from read alignment to DNA methylation analysis 93%
- SAMURAI: Shallow Analysis of copy nuMber alterations Using a Reproducible And Integrated bioinformatics pipeline 93%
- binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets 93%
Similar papers in this journal
- HQAlign: Aligning nanopore reads for SV detection using current-level modeling 94%
- HiC-TE: a computational pipeline for Hi-C data analysis shows a possible role of repeat family interactions in the genome 3D organization 94%
- OctopusV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.