Leveraging long-read assemblies and machine learning to enhance short-read transposable element detection and genotyping
Daigle, A. T.; Whitehouse, L. S.; Zhao, R.; Emerson, J. J.; Schrider, D. R.
Show abstract
Transposable elements (TEs) are parasitic genomic elements that are ubiquitous across the tree of life and play a crucial role in genome evolution. Advances in long-read sequencing have allowed highly accurate TE detection, though at a higher cost than short-read sequencing. Recent studies using long reads have shown that existing short-read TE detection methods perform inadequately when applied to real data. In this study, we use a machine learning approach (called TEforest) to discover and genotype TE insertions and deletions with short-read data by using TEs detected from long-read genome assemblies as training data. Our method first uses a highly sensitive algorithm to discover potential TE insertion or deletion sites in the genome, extracting relevant features from short-read alignments. To discriminate between true and false TE insertions, we train a random forest model with a labeled ground-truth dataset for which we have calculated the same set of short-read features. We conduct a comprehensive benchmark of TEforest and traditional TE detection methods using real data, finding that TEforest identifies more true positives and fewer false positives across datasets with different read lengths and coverages, while also accurately inferring genotypes and the precise breakpoints of insertions. By learning short-read signatures of TEs previously only discoverable using long reads, our approach bridges the gap between large-scale population genetic studies and the accuracy of long-read assemblies. This work provides a user-friendly tool to study the prevalence and phenotypic effects of TE insertions across the genome.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TT-Mars: Structural Variants Assessment Based on Haplotype-resolved Assemblies 95%
- Long-read detection of transposable element mobilization in the soma of hypomethylated Arabidopsis thaliana individuals 94%
- Benchmarking Transposable Element Annotation Methods for Creation of a Streamlined, Comprehensive Pipeline 94%
Similar papers in this journal
- TeloSearchLR: an algorithm to detect novel telomere repeat motifs using long sequencing reads 93%
- Minimizing detection bias of somatic mutations in a highly heterozygous oak genome 93%
- A new high-quality genome assembly and annotation for the threatened Florida Scrub-Jay (Aphelocoma coerulescens) 92%
Similar papers in this journal
- Earl Grey: a fully automated user-friendly transposable element annotation and analysis pipeline 97%
- Patterns of piRNA regulation in Drosophila revealed through transposable element clade inference 96%
- Dynamics and impacts of transposable element proliferation during the Drosophila nasuta species group radiation 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.