Back

DNA StairLoop: Achieving High Error-correcting and Parallel-processing Capabilities in DNA-based Data Storage

Yan, Z.; Qu, G.; Chen, X.; Zheng, G.; Wu, H.

2024-11-08 synthetic biology
10.1101/2024.11.07.622581 bioRxiv
Show abstract

DNA-based data storage is a promising solution to the challenges of large-scale data storage. However, the low throughput of the mainstream inkjet-based DNA synthesis method has hindered its widespread adoption. In contrast, high-throughput electrochemical synthesis provides higher throughput but with more nucleotide insertion, deletion, and substitution errors. Here, we propose an innovative coding scheme with high error correction capabilities, called DNA StairLoop. This coding scheme features a staircase interleaver and allows for component codes such as the convolutional code and the ow-density Parity-check code, allowing for flexible adaptation of the coding scheme. Both the row and the column decoders are soft input and soft output, enabling further improvement in data recovery accuracy through iterative decoding. The staircase interleaver facilitates extensive parallel decoding capabilities while effectively preserving parallelism across a multitude of nodes. In the in vitro experiment, DNA StairLoop successfully recovered the raw information in a staircase block with a nucleotide error rate of more than 8%. The simulations revealed that the DNA StairLoop can correct 10% nucleotide errors. Moreover, in parallel computing processing, the decoding time of our code continues to decrease dramatically.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.