Back

UniWave: A Waveform-Based Encoding Framework for Nucleic Acid Feature Extraction

Zhang, H.; Qi, Y.; Wang, L.

2026-01-13 bioinformatics
10.64898/2026.01.12.698567 bioRxiv
Show abstract

Nucleic acid sequence analysis constitutes a core research area in biomedical and health informatics, playing a critical role in infectious disease surveillance, epigenetic regulation, and genomic biomarker discovery. However, most existing sequence encoding methods rely on discrete representations, which are inadequate for capturing the intrinsic continuous structural properties of biological sequences. Inspired by waveform representations in modern physics, we propose UniWave, a novel encoding framework that converts discrete nucleotide sequences into biologically informative one-dimensional continuous waveforms through base mapping, windowed sinc interpolation, and wavelet-based downsampling. To enable more efficient spatiotemporal modeling, we design a learnable positional encoding module, WavePosition, which incorporates positional information to project the one-dimensional waveform into a two-dimensional continuous representation. In addition, we develop a lightweight dual-attention network, InceptionTime-ATT, to facilitate efficient multi-scale extraction of biological features. Experiments conducted on multiple representative genomic benchmark tasks demonstrate that UniWave consistently and significantly outperforms conventional encoding methods in viral genotype classification, cross-species enhancer identification, epigenetic modification site detection, and bacterial genome classification. Moreover, UniWave maintains excellent robustness under perturbation and noise stress tests, while its low-dimensional continuous representation further contributes to reduced model complexity. In summary, UniWave establishes a novel paradigm for biomedical nucleic acid sequenceencoding that is interpretable, noise-resilient, and computationally efficient, offering new avenues for uncovering disease-associated genomic patterns and advancing data-driven biomedical research.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.