REPrise: de novo interspersed repeat detection using inexact seeding
Takeda, A.; Nonaka, D.; Imazu, Y.; Fukunaga, T.; Hamada, M.
Show abstract
MotivationInterspersed repeats occupy a large part of many eukaryotic genomes, and thus their accurate annotation is essential for various genome analyses. Database-free de novo repeat detection approaches are powerful for annotating genomes that lack well-curated repeat databases. However, existing tools do not yet have sufficient repeat detection performance. ResultsIn this study, we developed REPrise, a de novo interspersed repeat detection software program based on a seed-and-extension method. Although the algorithm of REPrise is similar to that of RepeatScout, which is currently the de facto standard tool, we incorporated three unique techniques into REPrise: inexact seeding, affine gap scoring and loose masking. Analyses of rice and simulation genome datasets showed that REPrise outperformed RepeatScout in terms of sensitivity, especially when the repeat sequences contained many mutations. Furthermore, when applied to the complete human genome dataset T2T-CHM13, REPrise demonstrated the potential to detect novel repeat sequence families. AvailabilityThe source code of REPrise is freely available at https://github.com/hmdlab/REPrise. Repeat annotations predicted for the T2T genome using REPrise are also available at https://waseda.box.com/v/REPrise-data. Contactfukunaga@aoni.waseda.jp and mhamada@waseda.jp
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TandemMapper and TandemQUAST: mapping long reads and assessing/improving assembly quality in extra-long tandem repeats 95%
- HiC-TE: a computational pipeline for Hi-C data analysis shows a possible role of repeat family interactions in the genome 3D organization 95%
- The String Decomposition Problem and its Applications to Centromere Assembly 95%
Similar papers in this journal
- DANTE and DANTE_LTR: Lineage-centric annotation pipelines for long terminal repeat retrotransposons in plant genomes 95%
- Fast and memory-efficient mapping of short bisulfite sequencing reads using a two-letter alphabet 94%
- EASYstrata: An All-in-One Workflow for Genome Annotation and Genomic Divergence Analysis 92%
Similar papers in this journal
- The genome polishing tool POLCA makes fast and accurate corrections in genome assemblies 93%
- Identifying promoter sequence architectures via a chunking-based algorithm using non-negative matrix factorisation 93%
- Demonstrating the utility of flexible sequence queries against indexed short reads with FlexTyper 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.