Back

tinyRNA-seq: An optimized approach to sequencing tiny RNAs and primitive RNA genomes

Colville, B. W. F.; Zhao, J.; Hade, L.; Szostak, J. W.

2026-08-07 biochemistry
10.64898/2026.08.06.743385 bioRxiv
Show abstract

Very short RNAs play critical roles in modern biology, and are thought to have been crucial for genome replication during the origin of life. Next-generation sequencing is an essential tool for characterizing pools of small RNAs, but current library preparation methods suffer from strong size and sequence biases. Here we present tinyRNA-seq, an optimized library preparation method designed to minimize length- and sequence-dependent capture bias enabling the sequencing of RNA fragments as short as 2 nucleotides. We use degenerate adaptor regions to reduce ligation sequence bias and facilitate unique molecular identifier (UMI) installation. We benchmarked tinyRNA-seq against commercial kits using a model primordial RNA genome consisting of hundreds of defined oligonucleotides ranging from 2 to 12 nucleotides. tinyRNA-seq reproduced the input RNA distribution without the size and sequence bias of the commercial kits. tinyRNA-seq also enables the detection of de novo oligonucleotide generation, an important process for the origins of life. Applied to biologically derived small RNAs including miRNAs, piRNAs, and cityRNAs, tinyRNA-seq showed significantly lower capture bias and recovered a wider range of sequences than commercial kits. tinyRNA-seq may thus provide a more complete and quantitatively accurate representation of small RNAs from both biological and chemical sources. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=94 SRC="FIGDIR/small/743385v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@95ee64org.highwire.dtl.DTLVardef@155fb06org.highwire.dtl.DTLVardef@1d3665forg.highwire.dtl.DTLVardef@1e61404_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.