Benchmarking the coding strategies of non-coding mutations on sequence-based downstream tasks with machine learning
Liu, Z.; Bao, Y.; Li, W.; Li, W.; Lin, G. N.
10.1101/2025.01.01.631025 bioRxivShow abstract
Non-coding single nucleotide polymorphisms (SNPs) are key modulators of gene regulation and have been implicated in diverse complex traits and diseases. With the growing demand for accurate functional interpretation of non-coding variants, the choice of encoding strategies becomes critical in downstream predictive modeling. Despite recent advances, a systematic evaluation of encoding approaches tailored for non-coding SNPs remains lacking. To address this gap, we present a comprehensive benchmark that evaluates six representative encoding strategies, including categorical, semantic, and functional embeddings, across three quantitative trait loci (QTL)-related prediction tasks. The study encompasses nine machine learning and deep learning models and incorporates experimental controls and repeated trials to ensure robustness and reproducibility. We assess each strategy along multiple dimensions, such as interpretability, representation abundance, and computational efficiency. Rather than ranking individual methods, our analysis emphasizes the interaction between encoding strategies, model types, and preprocessing protocols, and highlights their collective influence on predictive performance. This work establishes a standardized framework for evaluating non-coding SNP representations and offers actionable guidance for selecting and optimizing prediction pipelines in regulatory genomics.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- WEVar: a novel statistical learning framework for predicting noncoding regulatory variants 96%
- A novel haplotype-based eQTL approach identifies genetic associations not detected through conventional SNP-based methods 94%
- Predicting 3D genome architecture directly from the nucleotide sequence with DNA-DDA 94%
Similar papers in this journal
- Evidence for the role of transcription factors in the co-transcriptional regulation of intron retention 95%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 94%
- COCOA: Coordinate covariation analysis of epigenetic heterogeneity 94%
Similar papers in this journal
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 95%
- SPREd: A simulation-supervised neural network tool for gene regulatory network reconstruction 95%
- Optimizer's dilemma: optimization strongly influences model selection in transcriptomic prediction 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.