Back

Polus: a Transformer-based Soft-decision Codec Enhancement Platform for DNA Storage

Ding, L.; Wang, K.; Zhang, H.; Xie, S.; Wang, J.; Liu, B.; Wang, G.; Liu, L.; Zhu, Z.

2025-12-18 bioinformatics
10.64898/2025.12.16.694663 bioRxiv
Show abstract

DNA storage offers exceptional information density and archival longevity, but is constrained by the complex, heterogeneous errors inherent to synthesis, storage, and sequencing. Conventional error-correction schemes often rely on excessive logical redundancy to mitigate these biochemical imperfections, thereby compromising storage efficiency. Here, we introduce Polus, a deep-learning-enabled platform that bridges the gap between biochemical constraints and digital reliability through soft-decision decoding. At its core is SeqFormer, a Transformer-based channel model that synergizes sequence context with quality signals to characterize platform-specific error profiles, generating calibrated per-base confidence scores. This mechanism transforms uncertain biochemical noise into informative "soft" erasures. In in silico benchmarks, Polus significantly enhances mainstream codecs: it reduces the sequencing coverage required for DNA Fountain by 38.9% --increasing effective physical density by [~]80%--and eliminates persistent indel-induced errors in the Yin-Yang codec. Furthermore, it enables a targeted resequencing strategy that achieves full recovery with 99.9% less overhead than brute-force deepening. To formalize these gains and address the lack of systematic benchmarking in the field, Polus establishes a standardized nine-metric evaluation framework that rigorously quantifies the trade-offs between reliability, density, and cost. This work provides a reproducible, quantitative foundation for next-generation, context-aware DNA storage systems.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.