Back

A ratio-based framework using Quartet reference materials for integrating long- and short-read RNA-seq

Chen, Q.; Guo, X.; Wang, D.; Zhao, J.; Xu, Y.; You, Y.; Mai, Y.; Duan, S.; Liu, Y.; Zhang, Y.; Li, X.; Chen, H.; Hou, W.; Yu, Y.; Dong, L.; Li, J.; Ritchie, M. E.; Zhang, R.; Shi, L.; Zheng, Y.

2025-09-17 genomics
10.1101/2025.09.15.676287 bioRxiv
Show abstract

Long-read RNA sequencing (lrRNA-seq) enables full-length transcript profiling but is confounded by technical batch effects that compromise quantification and prevent data integration across platforms, protocols, and laboratories. The lack of a transcriptome-wide biological ground truth has hindered objective benchmarking. To address these dual challenges, we leveraged certified Quartet reference materials to generate one of the largest multi-center lrRNA-seq resources to date: over one billion long reads from 144 libraries across four PacBio and Nanopore protocols in four independent laboratories. We first establish that ratio-based quantification against built-in reference samples effectively removes technical noise, revealing underlying biological signals. We then constructed the first ratio-based reference datasets for full-length transcripts-- comprising 10,218 isoforms and 6,032 alternative splicing (AS) events--and orthogonally validated them with RT-qPCR. Finally, a comprehensive benchmark using these ground truths reveals that a hybrid strategy integrating long- and short-read data (hybrid-seq) achieves the highest quantification accuracy for both isoforms and AS events. Our work provides a foundational framework and resource for evaluating lrRNA-seq technologies and accelerating the standardization of full-length transcriptomics for research and clinical applications.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.