Comprehensive Transcriptome Quality Assessment Using CATS: Reference-free and Reference-based Approaches
Bodulic, K.; Vlahovicek, K.
Show abstract
Accurate assessment of transcriptome assembly quality is critical to ensure the reliability of subsequent transcriptomic analyses. We present CATS (Comprehensive Assessment of Transcript Sequences), a tool offering both reference-free (CATS-rf) and reference-based (CATS-rb) transcriptome quality evaluation pipelines. CATS-rf maps RNA-seq reads back to the assembled transcripts and computes four interpretable scoring components that capture common assembly errors. CATS-rb assesses transcriptome completeness via alignment to a reference genome, supporting both annotation-free and annotation-based scoring. We benchmarked CATS on 672 transcriptomes from simulated and public RNA-seq data. CATS-rf outperformed existing tools in both transcript-level accuracy assessment and demonstrated high sensitivity to diverse assembly error types. CATS-rb produced robust transcriptome completeness estimates even without external annotation, with its scoring metrics strongly reflecting assembly quality. These results highlight CATS as an accurate, interpretable, and broadly applicable framework for evaluating transcriptome assemblies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RATTLE: Reference-free reconstruction and quantification of transcriptomes from Nanopore sequencing 96%
- SQANTI-SIM: a simulator of controlled transcript novelty for lrRNA-seq benchmark 96%
- Two-pass alignment using machine-learning-filtered splice junctions increases the accuracy of intron detection in long-read RNA sequencing 95%
Similar papers in this journal
- PCLIPtools: A Robust Framework for Identifying RNA-Protein Interaction Sites from PAR-CLIP experiments. 95%
- Shiba: A versatile computational method for systematic identification of differential RNA splicing across platforms 95%
- noisyR: Enhancing biological signal in sequencing datasets by characterising random technical noise 95%
Similar papers in this journal
- Foreign RNA spike-ins enable accurate allele-specific expression analysis at scale 93%
- OctopusV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis 93%
- Accurate Estimation of Molecular Counts from Amplicon Sequence Data with Unique Molecular Identifiers 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.