Back

Improved RNA homology detection and alignment by automatic iterative search in an expanded database

Singh, J.; Paliwal, K.; Singh, J.; Litfin, T.; Zhou, Y.

2022-10-07 bioinformatics
10.1101/2022.10.03.510702 bioRxiv
Show abstract

Unlike 20-letter-coded proteins, RNA homologous sequences are notoriously difficult to detect because their 4-letter-coded sequences can quickly lose their sequence identity. As a result, employing secondary structures has been found necessary to improve the sensitivity and the accuracy of homolog search. However, exact secondary structures often are not known. As a result, Rfam, the de facto gold-standard of RNA homologous families, has to rely on manual curation and experimental secondary structure if available. Here, we showed that using a combination of BLAST and iterative INFERNAL searches along with an expanded sequence database leads multiple sequence alignments (MSA) that are comparable to those provided by Rfam MSAs, according to secondary structure extracted from mutational coupling analysis and alignment accuracy when compared to structure alignment. The fully automatic tool (RNAcmap2) allows making homolog search, multiple sequence alignment, and mutational coupling analysis for any non-Rfam RNA sequences with Rfam-like performance.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.