Back

CUPID-seq enables highly multiplexed amplicon sequencing via combinatorial in-line dual indexing

Fu, B.; Porter, R. L.; Shi, H.; Ea, A. C.; Espeleta, A. M.; Ambat, A.; Relman, D. A.; Huang, K. C.; Xue, K. S.

2026-05-21 microbiology
10.64898/2026.05.20.726713 bioRxiv
Show abstract

Targeted amplicon sequencing is widely used to profile genetic variation in defined genomic regions. In microbial ecology, for example, amplicon sequencing of the 16S and 18S ribosomal RNA genes has been transformative for characterizing microbial communities. However, on high-capacity sequencing platforms with patterned flow cells, throughput is constrained by the requirement for unique dual indexes (UDIs), which increases primer costs and limits the number of samples that can be pooled per sequencing run. Here, we introduce CUPID-seq (Combinatorial, Unique, Phased, In-line Dual-indexed sequencing), a highly multiplexed amplicon-sequencing strategy that increases scalability through combinatorial indexing across two rounds of PCR. CUPID-seq introduces phased, in-line UDIs during Round 1 gene-specific amplification, enabling multiple samples to share the same Illumina UDI during Round 2 PCR while remaining uniquely identifiable. This design reduces upfront costs by up to 85% and reduces library preparation time and reagent use by up to 40%. We develop and validate CUPID-seq primers targeting the 16S V4 region and provide a computational workflow for demultiplexing in-line indexes. Although optimized here for 16S-based profiling, CUPID-seq can be readily adapted to other user-defined amplicons. By reducing cost and increasing multiplexing capacity, CUPID-seq enables users to leverage high-throughput sequencing platforms more effectively across diverse biological contexts.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.