Back

Designing DNA With Tunable Regulatory Activity Using Score-Entropy Discrete Diffusion

Sarkar, A.; Kang, Y.; Somia, N.; Mantilla, P.; Zhou, J. L.; Nagai, M.; Tang, Z.; Zhao, C.; Koo, P.

2025-05-23 genomics
10.1101/2024.05.23.595630 bioRxiv
Show abstract

Designing regulatory DNA sequences with precise, cell-type-specific activity is critical for applications in medicine and biotechnology, but remains challenging due to the vast combinatorial space and complex regulatory grammar governing gene expression. Recent deep generative models--including genomic language models and diffusion-based approaches--offer new tools for sequence design, yet lack systematic evaluation frameworks to assess the biological and functional fidelity of generated sequences. Here, we introduce a comprehensive computational framework for evaluating generated sequences based on their functional activity, sequence similarity, and regulatory motif composition relative to natural regulatory DNA. We further present DNA Discrete Diffusion (D3), a score-entropy discrete diffusion model for conditional generation of regulatory sequences. Benchmarking D3 on multiple functional genomics datasets, we find that D3 produces sequences nearly indistinguishable from natural DNA under our evaluation metrics. Unlike previous diffusion models, which often fail to capture the nuanced combinatorial patterns of regulatory elements, D3 effectively recapitulates cell-type-specific activity and motif organization. We also show that D3 learns informative representations even in the absence of conditioning labels, outperforming genomic language models and supervised models trained on naive one-hot encodings. D3 maintains strong performance in low-data regimes and enhances downstream supervised models when its generated sequences are used for data augmentation. Together, our work advances generative design of regulatory DNA and establishes comprehensive evaluation methods to ensure biological fidelity.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.