DNA Sequence Trace Reconstruction Using Deep Learning
Cao, B.; Xie, L.; Liu, Z.; Li, X.; Wang, B.; Zhou, S.; Zheng, P.; Zhang, Q.
Show abstract
Deciphering DNA sequences is fundamental to unlocking the mysteries of life, but the high dimensionality and complexity of biological sequence data significantly hinder knowledge discovery. In particular, the challenges of sequence length, repetitive regions, and structural complexity make it difficult to directly reconstruct complete DNA sequences from raw data. Therefore, this paper proposes a DNA sequence trace reconstruction model, DNARetrace, which performs preprocessing and dataset construction, and then employs a Bidirectional Fourier-Kolmogorov-Arnold Network (Bi-FKGAT), using an extremely unbalanced loss function for link prediction, so as to reconstruct the original DNA sequence. In multi-angle experiments using both simulated and real data, DNARetrace successfully reconstructs DNA sequence traces across large-scale datasets derived from various DNA sequencing methods, overcoming the bias of current approaches toward specific sequencing platforms, and achieves competitive outcomes in DNA storage and genomics downstream tasks. We further validated the expandability of the proposed methods in DNA sequence classification and metagenomic binning tasks. In summary, DNARetrace is compatible with various sequencing scenarios; it reduces the difficulty of discovering novelty knowledge directly from high-complexity raw data, and it provides a reusable tool to accelerate DNA sequence processing and applications.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- deSALT: fast and accurate long transcriptomic read alignment with de Bruijn graph-based index 97%
- Sequences Dimensionality-Reduction by K-mer Substring Space Sampling Enables Effective Resemblance- and Containment-Analysis for Large-Scale omics-data 96%
- BERMUDA: A novel deep transfer learning method for single-cell RNA sequencing batch correction reveals hidden high-resolution cellular subtypes 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.