Back

Screening metatranscriptomes for ultrastable RNA secondary structures reveals hidden bacteriophages and novel capsid nanomaterials

Villarreal, D. A.; Makasarashvili, N.; Kapoor, A.; Root, M.; Campbell, M.; Gibson, S.; Schiveley, C.; Rastandeh, A.; Baker, S.; Subramanian, S.; Neri, U.; Mills, C. E.; McNair, K.; Segall, A. M.; Gophna, U.; Parent, K. N.; Garmann, R. F.

2026-04-07 biochemistry
10.64898/2026.04.05.716407 bioRxiv
Show abstract

Metatranscriptomics has transformed our view of RNA bacteriophage diversity, revealing vast numbers of single-stranded RNA (ssRNA) phages whose protein capsids can be engineered for biotechnology applications. However, many ssRNA phages remain hidden from current detection methods, which require protein-level similarity to known phages. Here we show that RNA structure provides an additional signal for the detection of ssRNA phages in metatranscriptomes, including hidden phages missed by prior protein-based methods. By computationally folding each contig and screening for exceptionally stable RNA secondary structures, we find evidence of thousands of previously unrecognized phages encoding novel coat proteins. We express a library of 12,000 such coat proteins in E. coli and find that most assemble into nuclease-resistant capsids. We determine the 3D structure of one such capsid by cryo-electron microscopy and demonstrate that it can be disassembled and reassembled in vitro to package heterologous RNA--a key step toward repurposing these particles as RNA delivery vehicles. We compile the newly discovered ssRNA phages with previously known ones into a database that contains sequence and structural information for over 460,000 unique RNA molecules and over 100,000 distinct coat proteins, providing a comprehensive resource for microbiology and nanomaterials research.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.