Back

Intrinsic molecular identifiers enable robust molecular counting in single-cell sequencing

Fontanez, K.; Agam, Y.; Bevans, S.; May-Zhang, A.; Hayford, C.; Xue, Y.; Ishibashi, J. S. A.; Komuhendo, R.; Yoder, L.; Hettige, P.; Rickner, H.; Zhang, J. Q.; D'Amato, C.; Smithers, T.; Osman, A.; Calkins, S.; Rahman, M.; Mutafopulos, K. S.; Kiani, S.; Meltzer, R. H.

2024-10-05 genomics
10.1101/2024.10.04.616561 bioRxiv
Show abstract

Particle-templated instant partition sequencing (PIPseq), is an emerging approach for massively scalable single-cell gene expression studies that does not require complex instrumentation or expensive consumables. We present PIPseqTM V, a novel implementation of the PIPseq workflow with significant improvements in assay performance and sensitivity compared to prior methods. Among the innovations driving the improved performance in PIPseq V is a new approach for transcript counting using Intrinsic Molecular Identifiers (IMIs) from the captured transcript sequence, eliminating the need for traditional Unique Molecular Identifiers (UMIs). IMIs are the starting positions for molecules generated by random fragmentation after limited cycle PCR, which can be used to uniquely identify individual transcript copies of genes expressed within individual cells. To correct for experimental variation in IMI generation, we have developed a dynamic correction strategy that can adapt to different sample cell input, transcript expression, and sequencing depth without the need for external benchmarking or controls. Our results demonstrate that PIPseq V with IMI-based analysis provides biological information comparable to established UMI-based approaches while avoiding UMI-associated biases. Dynamic correction offers a robust and data-driven analysis strategy for scRNAseq.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.