Experimental Confirmation of Multiple Co-Existent DNA Secondary Structures using Low-Yield Bisulfite Sequencing
Li, J.; Bae, J.; Yordanov, B.; Wang, M. X.; Gonzalez, J.; Philips, A.; Zhang, D. Y.
Show abstract
Predicting DNA secondary structures is critical to a broad range of applications involving single-stranded DNA (ssDNA), yet remains an open problem. Existing prediction models are limited by insufficient experimental data, due to a lack of high-throughput methods to study DNA structures, in contrast to RNA structures. Here, we present a method for profiling DNA secondary structures using multiplexed low-yield bisulfite sequencing (MLB-seq), which examines the chemical accessibility of cytosines in thousands of different oligonucleotides. By establishing a probability-based model to evaluate the consensus probability between MLB-seq data and structures proposed using NUPACK software, we identified the secondary structures of individual ssDNA molecules and estimated the distribution of multiple secondary structures in solution. We studied the structures of 1,057 human genome subsequences and experimentally confirmed that 84% adopted two or more structures. MLB-seq thus enables high-throughput ssDNA structure profiling and will benefit the design of probes, primers, aptamers, and genetic regulators.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Training Data Diversity Enhances the Basecalling of Novel RNA Modification-Induced Nanopore Sequencing Readouts 97%
- High-Throughput DNA melt measurements enable improved models of DNA folding thermodynamics 96%
- Demultiplexing and barcode-specific adaptive sampling for nanopore direct RNA sequencing 96%
Similar papers in this journal
- Quantification of Cas9 binding and cleavage across diverse guide sequences maps landscapes of target engagement 93%
- scAllele: a versatile tool for the detection and analysis of variants in scRNA-seq 93%
- DNB-Based On-Chip Motif Finding (DocMF): a High-Throughput Method to Profile Different Types of Protein-DNA Interactions 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.