SEEDS: Simulating Emergence of Errors in DNA Storage
Naznin, M. F. S.
Show abstract
BackgroundDNA storage is a nonvolatile memory technology for storing data as synthetic DNA strings which offers unprecedented storage density and durability. Yet, the application of DNA as a practical digital information storage medium remains an enigma, since this is extremely expensive and it takes a substantial amount of time to encode and decode data to/from synthetic DNA. More importantly, various phases of DNA storage pipeline (e.g., synthesis, sequencing, etc.) are error prone. Furthermore, DNA is subject to decay over time and the reliability of the synthetic DNA depends on various aspects, including preservation medium and temperature. To allow for the perfect storage and recovery of the information and thereby making it competitive with the existing flash or tape based technologies, advanced error protection schemes are necessary. However, evaluating and comparing various DNA storage technologies and error correcting codes under realistic model conditions - comprising a wide array of synthesis medium, sequencing technologies, temperature and duration - is prohibitively time consuming and expensive. ResultsIn this study, we present SEEDS, an error model based simulator to mimic the process of accumulating errors at different phases of DNA storage. SEEDS is the first known simulator which incorporates various empirically derived statistical (or stochastic?) error models, mimicking the generation and propagation of different types of errors at various phases in DNA storage. It was assessed for its validity against the data from a number of published wet-lab experiments. ConclusionsSEEDS is easy to use and offers flexible and comprehensive parameter settings to mimic the error models in DNA storage. Validation against in vitro experimental results suggests its promise for emulating the stochastic models of error generation and propagation in DNA storage. SEEDS is available as a web interface with a server side application, along with portable cross-platform native applications (available at givethelink).
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Beam search decoder for enhancing sequence decoding speed in single-molecule peptide sequencing data 95%
- MCell4 with BioNetGen: A Monte Carlo Simulator of Rule-Based Reaction-Diffusion Systems with Python Interface 94%
- A mechanistic model of the BLADE platform predicts performance characteristics of 256 different synthetic DNA recombination circuits 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.