mclUMI: Markov clustering of unique molecular identifiers enables dynamic removal of PCR duplicates
Sun, J.; Li, S.; Canzar, S.; Cribbs, A. P.
Show abstract
Molecular quantification in high-throughput sequencing experiments relies on accurate identification and removal of polymerase chain reaction (PCR) duplicates. The use of Unique Molecular Identifiers (UMIs) in sequencing protocols has become a standard approach for distinguishing molecular identities. However, PCR artefacts and sequencing errors in UMIs present a significant challenge for effective UMI collapsing and accurate molecular counting. Current computational strategies for UMI collapsing often exhibit limited flexibility, providing invariable deduplicated counts that inadequately adapt to varying experimental conditions. To address these limitations, we developed mclUMI, a tool employing the Markov clustering algorithm to accurately identify original UMIs and eliminate PCR duplicates. Unlike conventional methods, mclUMI automates the detection of independent communities within UMI graphs by dynamically fine-tuning inflation and expansion parameters, enabling context-dependent merging of UMIs based on their connectivity patterns. Through in silico experiments, we demonstrate that mclUMI generates dynamically adaptable deduplication outcomes tailored to diverse experimental scenarios, particularly best-performing under high sequencing error rates. By integrating connectivity-driven clustering, mclUMI enhances the accuracy of molecular counting in noisy sequencing environments, addressing the rigidity of current UMI deduplication frameworks.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genotype-free demultiplexing of pooled single-cell RNA-Seq 97%
- deSALT: fast and accurate long transcriptomic read alignment with de Bruijn graph-based index 96%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 96%
Similar papers in this journal
- XCVATR: Detection and Characterization of Variant Impact on the Embeddings of Single -Cell and Bulk RNA-Sequencing Samples 95%
- MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores 95%
- DNAscent v2: Detecting Replication Forks in Nanopore Sequencing Data with Deep Learning 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.