DNA-SaM, a robust system for large-scale data storage
Huang, X.; Wang, Y.; Xu, J.; Nie, Z.; Huang, J.; Wu, Y.; Qin, Z.; Dai, J.; Wang, Y.
Show abstract
DNA data storage offers a viable strategy to address the impending data explosion. Early attempts to harness DNA as a storage medium have encountered scalability limitations, largely due to the complexity of codec algorithms, the generation of biochemically harmful sequences and lack of a robust architecture. We present "DNA-SaM", a novel system designed for DNA data storage, which achieves linear computational complexity and strict bio-constraint adherence, ensuring high coding efficiency and fidelity. It encoded data at speeds surpassing classic systems by over 2 orders of magnitude, with this superiority changes across various encoding algorithms. Importantly, DNA-SaM effectively eliminates any sequence that could be deleterious to in vitro and in vivo biochemical processes, including homopolymer runs, tandem repeat motifs, and potential promoter sequences, etc. It also involves an advanced DNA data storage architecture that incorporates a two-tiered indexing system and a novel "storage unit" distribution paradigm for large-scale data storage. It is further validated by practical data storage both in vitro and in vivo with a 100% success rate. Our system is capable of storing data over 1039 PB, which marks a critical advancement in the scalability of DNA-based data storage.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Gra-CRC-miRTar: The pre-trained nucleotide-to-graph neural networks to identify potential miRNA targets in colorectal cancer 92%
- CertPrime: a new oligonucleotide design tool for gene synthesis 92%
- DeepCORE: An interpretable multi-view deep neural network model to detect co-operative regulatory elements 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.