Back

Directional purine pyrimidine asymmetry in CRISPR repeats facilitates spacer acquisition

Barik, S.; Sahu, P.; Ghosh, K.; Subramanian, H.

2026-07-21 genomics
10.64898/2026.07.20.739571 bioRxiv
Show abstract

Spacer acquisition is the primary and essential step of CRISPR-Cas adaptive immunity in most prokaryotes and occurs preferentially at the leader-repeat junction of the CRISPR array. Despite the conservation of the adaptation machinery, spacer acquisition efficiencies and site-specificity vary markedly across CRISPR-Cas systems, suggesting that the local sequence architecture of the repeat and leader-repeat junction may contribute to the variation in the adaptation efficiency. To investigate this possibility, we systematically analyzed the direct repeat sequences and the terminal base pairs at the 3' end of leader regions adjacent to the integration site. We identify a conserved asymmetric purine-pyrimidine (RY) distribution within direct repeats, characterized by a 5' pyrimidine-rich half and a 3' purine-rich half, together with purine enrichment at the 3' end of leader sequences. Building on these observations, we develop a mechanistic model for spacer acquisition by demonstrating the importance of the sequence motif of the first direct repeat and the leader-repeat junction, using the sequence-dependent asymmetric cooperativity model that captures DNA unzipping kinetics. According to our model, sequences with such directional RY-asymmetric nucleotide distribution as direct repeats have high unzipping propensity and thus can promote efficient spacer integration. Consistent with this mechanism, experimental data from previous studies show that repeats with stronger asymmetry and appropriately positioned pyrimidine-to-purine transition sites are associated with larger spacer counts within CRISPR arrays. These findings reveal a conserved architectural feature of CRISPR repeats and provide a mechanistic model connecting DNA sequence organization to efficiency.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.