Back

Rfam 15: RNA families database in 2025

Ontiveros-Palacios, N.; Cooke, E.; Nawrocki, E. P.; Triebel, S.; Marz, M.; Rivas, E.; Griffiths-Jones, S.; Petrov, A. I.; Bateman, A.; Sweeney, B.

2024-09-24 genomics
10.1101/2024.09.23.614430 bioRxiv
Show abstract

The Rfam database, a widely-used repository of non-coding RNA (ncRNA) families, has undergone significant updates in release 15.0. This paper introduces major improvements, including the expansion of Rfamseq to 26, 106 genomes, a 76% increase, incorporating the latest UniProt reference proteomes and additional viral genomes. Sixty-five RNA families were enhanced using experimentally determined 3D structures, improving the accuracy of consensus secondary structures and annotations. R-scape covariation analysis was used to refine structural predictions in 26 families. Gene Ontology and Sequence Ontology annotations were comprehensively updated, increasing GO term coverage to 75% of families. The release adds 14 new Hepatitis C Virus RNA families and completes microRNA family synchronisation with miRBase, resulting in 1, 603 microRNA families. New data types, including FULL alignments, have been implemented. Integration with APICURON for improved curator attribution and multiple website enhancements further improve user experience. These updates significantly expand Rfams coverage and improve annotation quality, reinforcing its critical role in RNA research, genome annotation, and the development of machine learning models. Rfam is freely available at https://rfam.org. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=112 SRC="FIGDIR/small/614430v1_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@9b1898org.highwire.dtl.DTLVardef@6bc030org.highwire.dtl.DTLVardef@16b789org.highwire.dtl.DTLVardef@16b8577_HPS_FORMAT_FIGEXP M_FIG Rfam has undergone a major update with the release of 15.0. We have increased the number of genomes in our sequence database Rfamseq by 75%, completed the synchronisation with miRBase and improved 65 families using 3D structures. C_FIG

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.