Back

Discovery of Inverted Repeats and Palindromes in Viral Genomes

Sen, M.; Rivera, G. M.; Gao, J.; Shtrahman, M.; Ganapathiraju, M.

2025-11-12 bioinformatics
10.1101/2025.11.10.687097 bioRxiv
Show abstract

An inverted repeat (IR) in DNA is a sequence of nucleotides that is followed by its complementary bases but in reverse order, occurring on the same strand (e.g., TCACCGCGGTGA). If the two complementary sequences occur one after the other without other bases between them, they are referred to as DNA palindromes. IRs could form hairpin and cruciform secondary structures, which endanger genomic stability. They are found to be prevalent in viral DNA at origins of replication, and they play a crucial role in various biological processes including gene silencing, duplication, and genomic evolution. IRs have been less explored, which stems from the scarcity of sequence analysis tools allowing accurate detection on large viral genome data. Here, using the Biological Language Modeling Toolkit (BLMT), we analyzed 14 thousand viral genomes for occurrences of IRs, resulting in the identification of over 19 million IRs longer than 20 bases, including 134 IRs that are 2000 bases long, and around 1,300 IRs per virus.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.