Back

REPrise: de novo interspersed repeat detection using inexact seeding

Takeda, A.; Nonaka, D.; Imazu, Y.; Fukunaga, T.; Hamada, M.

2024-01-24 bioinformatics
10.1101/2024.01.21.576581 bioRxiv
Show abstract

MotivationInterspersed repeats occupy a large part of many eukaryotic genomes, and thus their accurate annotation is essential for various genome analyses. Database-free de novo repeat detection approaches are powerful for annotating genomes that lack well-curated repeat databases. However, existing tools do not yet have sufficient repeat detection performance. ResultsIn this study, we developed REPrise, a de novo interspersed repeat detection software program based on a seed-and-extension method. Although the algorithm of REPrise is similar to that of RepeatScout, which is currently the de facto standard tool, we incorporated three unique techniques into REPrise: inexact seeding, affine gap scoring and loose masking. Analyses of rice and simulation genome datasets showed that REPrise outperformed RepeatScout in terms of sensitivity, especially when the repeat sequences contained many mutations. Furthermore, when applied to the complete human genome dataset T2T-CHM13, REPrise demonstrated the potential to detect novel repeat sequence families. AvailabilityThe source code of REPrise is freely available at https://github.com/hmdlab/REPrise. Repeat annotations predicted for the T2T genome using REPrise are also available at https://waseda.box.com/v/REPrise-data. Contactfukunaga@aoni.waseda.jp and mhamada@waseda.jp

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.