Back

A systematic benchmark of bioinformatics methods for single-cell and spatial RNA-seq Nanopore long-read data

Hamraoui, A.; Onfroy, A.; Senamaud-Beaufort, C.; Coulpier, F.; Lemoine, S.; Jourdren, L.; Thomas-Chollier, M.

2025-07-25 bioinformatics
10.1101/2025.07.21.665920 bioRxiv
Show abstract

Alternative splicing plays a crucial role in transcriptomic complexity, yet remains difficult to resolve at the single-cell level due to the limitations of short-read technologies. Coupling single-cell with long-read sequencing offers full-length transcript coverage, enabling more accurate isoform detection. Multiple specialized computational tools tailored for single-cell and spatial long-read transcriptomics have been developed, with diverse strategies. To compare the effectiveness of these approaches, we generated paired short-read and Nanopore long-read single-cell datasets, tailored for benchmarking bioinformatics tools. We evaluated ten state-of-the-art methods, spanning four analytical dimensions: barcodes and UMI detection, demultiplexing and UMI clustering, gene-level expression profiling, and isoform detection and quantification. Using real and simulated datasets across different protocols, sequencing depths and chemistries, we assessed the accuracy, robustness, and scalability of each tool. Our results revealed method-specific trade-offs, and highlight the importance of sequencing quality and UMI correction strategies. This benchmark provides a practical resource for optimizing isoform analysis and accurate gene expression profiling in single-cell and spatial transcriptomics using long-read sequencing. Our benchmarking workflow is designed to be reusable, thereby enabling method developers to compare their own approaches against the set of reference methods evaluated in this work.

Published in NAR Genomics and Bioinformatics (predicted rank #5) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.