Back

siDiff: De Novo siRNA Design via Efficacy-Guided Discrete Masked Diffusion

Yue, Z.; Zhang, H.; Gao, X.; Shu, S.; Lai, L.

2026-08-11 bioinformatics
10.64898/2026.08.11.744092 bioRxiv
Show abstract

Target-conditioned de novo siRNA design requires capturing both target mRNA binding context and internal siRNA sequence-activity rules. Traditional computational pipelines predominantly rely on discriminative prediction models that rank pre-filtered candidate pools. However, these approaches struggle with generalization on novel target genes and often overlook potent candidates due to dataset size constraints and motif overfitting. To reconcile generation precision with sequence diversity, we propose siDiff, an efficacy-guided discrete diffusion framework for target-conditioned de novo siRNA design. siDiff pairs a discrete-masked diffusion transformer--which models the underlying sequence distribution over functional duplexes--with a mask-robust efficacy guidance model. During inference, we introduce a biology-aware, three-stage sampling mechanism that performs structural candidate filtering, dynamic unmasking guidance, and cluster-aware redundancy mitigation. This dual mechanism enables the diffusion process to explore broad sequence spaces while the efficacy model prevents distributional shift toward non-functional candidates. Extensive experiments across four datasets, including the public Takayuki benchmark and three curated patent datasets, demonstrate that siDiff significantly outperforms state-of-the-art discriminative baselines and discrete diffusion models, achieving relative hit-rate improvements of over 30% and demonstrating superior generalization on out-of-distribution gene targets. The source code and related materials are available at https://github.com/cybericha/siDiff.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.