Back

Benchmarking long-read RNA-sequencing technologies with LongBench: a cross-platform reference dataset profiling cancer cell lines with bulk and single-cell approaches

You, Y.; Solano, A.; Lancaster, J.; David, M.; Wang, C.; Su, S.; Zeglinski, K.; Ghamsari, R.; Chauhan, M.; Gleeson, J.; Prawer, Y. D. J.; Ng, J.; Dubois, B.; Cleynen, I.; Asselin-Labat, M.-L.; Sutherland, K. D.; Clark, M. B.; Gouil, Q.; Ritchie, M. E.

2025-09-12 genomics
10.1101/2025.09.11.675724 bioRxiv
Show abstract

Long-read RNA sequencing enables full-length transcript profiling and improved isoform resolution, but variable platforms and evolving chemistries demand careful benchmarking for reliable application. We present LongBench, a matched, multi-platform reference dataset spanning bulk, single-cell, and single-nucleus transcriptomics across eight human lung cancer cell lines with synthetic spike-in controls. LongBench incorporates three state-of-the-art long-read protocols alongside Illumina short reads: Oxford Nanopore Technologies (ONT) PCR-cDNA, ONT direct RNA, and PacBio Kinnex. We systematically evaluate transcript capture, quantification accuracy, differential expression, isoform usage, variant detection, and allele-specific analyses. Our results show high concordance in gene-level differential analyses across protocols, but reduced consistency for transcript-level and isoform analyses due to lengthand platform-dependent biases. Single-cell long-read data are highly concordant with bulk for high-confidence features, though single-nuclei data show reduced feature detection. LongBench provides one of the largest publicly available long-read benchmarking resources, enabling rigorous cross-platform evaluation and guiding technology selection for transcriptomic research.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.