Transcriptome Universal Single-isoform COntrol: A Framework for Evaluating Transcriptome reconstruction Quality
Liu, T.; Paniagua, A.; Jetzinger, F.; Ferrandez-Peral, L.; Frankish, A.; Conesa, A.
Show abstract
Long-read sequencing (LRS) platforms, such as Oxford Nanopore and Pacific Biosciences, enable comprehensive transcriptome analysis but face challenges such as sequencing errors, sample quality variability, and library preparation biases. Current benchmarking approaches address these issues insufficiently: BUSCO assesses transcriptome completeness using conserved single-copy orthologs but can misinterpret alternative splicing as gene duplications, while spike-ins (SIRVs, ERCCs) oversimplify real- sample complexity, neglecting RNA degradation and RNA extraction artifacts, thus inflating performance metrics. Simulation algorithms are limited to recapitulate this complexity. To overcome these limitations, we introduce the Transcriptome Universal Single-isoform Control (TUSCO), a curated internal reference set of genes lacking alternative isoforms. TUSCO evaluates precision by identifying transcripts deviating from reference annotations and assesses sensitivity by verifying detection completeness in human and mouse samples. Masking TUSCO transcripts--and optionally inserting decoy splice variants--creates a novel- isoform challenge that assesses recovery of the true, now-unannotated isoforms. Our validation demonstrates that TUSCO provides accurate and reliable benchmarking without external controls, significantly improving quality control standards for transcriptome reconstruction using LRS.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CapTrap-Seq: A platform-agnostic and quantitative approach for high-fidelity full-length RNA transcript sequencing 98%
- A spatial long-read approach at near-single-cell resolution reveals developmental regulation of splicing and polyadenylation sites in distinct cortical layers and cell types. 97%
- ORF Capture-Seq: a versatile method for targeted identification of full-length isoforms 96%
Similar papers in this journal
- Systematic assessment of long-read RNA-seq methods for transcript identification and quantification 97%
- SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms 97%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.