Back

Graph Attention Neural Networks Reveal TnsC Filament Assembly in a CRISPR-Associated Transposon

Pindi, C.; Ahsan, M.; Sinha, S.; Palermo, G.

2025-06-17 biophysics
10.1101/2025.06.17.659969 bioRxiv
Show abstract

CRISPR-associated transposons (CAST) enable programmable, RNA-guided DNA integration, marking a transformative advancement in genome engineering. A central player in the type V-K CAST system is the AAA+ ATPase TnsC, which assembles into helical filaments on double-stranded DNA (dsDNA) to orchestrate target site recognition and transposition. Despite its essential role, the molecular mechanisms underlying TnsC filament nucleation and elongation remain poorly understood. Here, multiple-microsecond and free energy simulations are combined with deep learning-based Graph Attention Network (GAT) models to elucidate the mechanistic principles of TnsC filament formation and growth. Our findings reveal that ATP binding promotes TnsC nucleation by inducing DNA remodelling and stabilizing key protein-DNA interactions, particularly through conserved residues in the initiator-specific motif (ISM). Furthermore, GNN-based attention analyses identify a directional bias in filament elongation in the 5'[->]3' direction and uncover a dynamic compensation mechanism between incoming and bound monomers that facilitate directional growth along dsDNA. By leveraging deep learning-based graph representations, our GAT model provides interpretable mechanistic insights from complex molecular simulations and is readily adaptable to a wide range of biological systems. Altogether, these findings establish a mechanistic framework for TnsC filament dynamics and directional elongation, advancing the rational design of CAST systems with enhanced precision and efficiency.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.