Back

Tresor: An integrated platform for simulating transcriptomic reads with realistic PCR error representation across various RNA sequencing technologies

Sun, J.; Cribbs, A. P.

2025-03-17 bioinformatics
10.1101/2025.03.15.643015 bioRxiv
Show abstract

The rapid advancement of high-throughput sequencing technologies has spurred the development of numerous computational tools designed to identify gene expression patterns from growing datasets at both bulk and single-cell sequencing levels. The recent advent of longread sequencing technologies has further accelerated the availability and refinement of these tools. The lack of ground-truth labels and annotations in sequencing data presents a significant challenge for evaluating the efficacy of analytical tools. To address this, we developed Tresor, an integrated platform for simulating both short and long reads at bulk and single-cell levels. We devised a tree-based algorithm to significantly accelerate in silico experiments at high PCR cycles. Tresor allows for customising sequencing libraries with highly modular and flexible read structures, facilitating the verification of sequencing-related biological discoveries. This tool also includes features that introduce substitution, insertion, and deletion errors at various stages of library preparation, PCR amplification, and sequencing, enhancing its applicability in diverse experimental conditions and simulating real world conditions. Our results demonstrate that, upon removal of PCR duplicates, cell type-specific gene expression profiles derived from our simulated reads highly resemble reference data. We envisage that Tresor will provide valuable insights into a broader range of transcriptomics analyses and support the development of more effective algorithms for read alignment and UMI deduplication.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.