Back

Quantitative Full-length transcriptome analysis by nanopore sequencing with Error-Aware UMI mapping

Li, Q.; Li, D.; Sun, C.; Song, G.; Huang, Y.; Lou, J.

2026-01-14 bioinformatics
10.64898/2026.01.14.699429 bioRxiv
Show abstract

Comprehensive transcriptome profiling is essential for understanding RNA diversity and regulation, yet accurate identification and quantification of full-length transcript isoforms remain challenging with short-read sequencing technologies. Nanopore sequencing enables direct sequencing of long cDNA molecules and thus offers a powerful solution for full-length transcriptome analysis, but its application to quantitative transcriptomics is limited by PCR amplification bias and the difficulty of unique molecular identifier (UMI) recognition under high sequencing error rates. Here, we developed UMImap, a dedicated pipeline for robust UMI identification, error correction, and deduplication in nanopore data. By integrating transcript-aware UMI correction with long-read isoform assembly, UMImap substantially improves UMI recognition accuracy compared with existing methods and effectively mitigates PCR-induced duplication bias. Using this framework, we identified tens of thousands of full-length transcript isoforms, including a large fraction of previously unannotated isoforms that are significantly longer than reference annotations. Quantitative analyses demonstrate the reliability of UMImap for transcript-level quantification. Functional and pathway enrichment analyses of highly expressed novel isoforms revealed coherent and biologically meaningful patterns, including strong enrichment in RNA processing, splicing, and translation pathways. Our results establish UMImap as an effective solution for UMI-based quantification in nanopore full-length transcriptome sequencing and highlight the potential of long-read sequencing to simultaneously achieve accurate isoform discovery and expression analysis in complex transcriptomes.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.