UniWave: A Waveform-Based Encoding Framework for Nucleic Acid Feature Extraction
Zhang, H.; Qi, Y.; Wang, L.
Show abstract
Nucleic acid sequence analysis constitutes a core research area in biomedical and health informatics, playing a critical role in infectious disease surveillance, epigenetic regulation, and genomic biomarker discovery. However, most existing sequence encoding methods rely on discrete representations, which are inadequate for capturing the intrinsic continuous structural properties of biological sequences. Inspired by waveform representations in modern physics, we propose UniWave, a novel encoding framework that converts discrete nucleotide sequences into biologically informative one-dimensional continuous waveforms through base mapping, windowed sinc interpolation, and wavelet-based downsampling. To enable more efficient spatiotemporal modeling, we design a learnable positional encoding module, WavePosition, which incorporates positional information to project the one-dimensional waveform into a two-dimensional continuous representation. In addition, we develop a lightweight dual-attention network, InceptionTime-ATT, to facilitate efficient multi-scale extraction of biological features. Experiments conducted on multiple representative genomic benchmark tasks demonstrate that UniWave consistently and significantly outperforms conventional encoding methods in viral genotype classification, cross-species enhancer identification, epigenetic modification site detection, and bacterial genome classification. Moreover, UniWave maintains excellent robustness under perturbation and noise stress tests, while its low-dimensional continuous representation further contributes to reduced model complexity. In summary, UniWave establishes a novel paradigm for biomedical nucleic acid sequenceencoding that is interpretable, noise-resilient, and computationally efficient, offering new avenues for uncovering disease-associated genomic patterns and advancing data-driven biomedical research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 94%
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 94%
- Learning, Visualizing and Exploring 16S rRNA Structure Using an Attention-based Deep Neural Network 94%
Similar papers in this journal
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 94%
- Phylogenetic-informed graph deep learning to classify dynamic transmission clusters in infectious disease epidemics 94%
- Towards Computing Attributions for DimensionalityReduction Techniques 94%
Similar papers in this journal
- Learning interpretable representations of single-cell multi-omics data with multi-output Gaussian Processes 94%
- Integrating experimental feedback improves generative models for biological sequences 93%
- OmicsFootPrint: a framework to integrate and interpret multi-omics data using circular images and deep neural networks 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.