A comprehensive benchmark and guide for sequence-function interpretable deep learning models in genomics
Sun, C.; Sun, Y.; Xu, K.; He, Z.; Li, H.; Li, Y.; Yu, Z.; wang, Y.; Lin, X.; Xu, X.; Hu, P.; Bo, X.; Liao, M.; Chen, H.
Show abstract
The development of sequence-based deep learning methods has greatly increased our understanding of how sequence determines function. In parallel, numerous interpretable algorithms have been developed to address complex tasks, such as elucidating sequence regulatory syntax and analyzing non-coding variants from trained models. However, few studies have systematically compared and evaluated the performance and interpretability of these algorithms. Here, we introduce a comprehensive benchmark framework for evaluating sequence-to-function models. We systematically evaluated multiple models and DNA language foundation models using 369 ATAC-seq datasets, employing diverse training strategies and evaluation metrics to uncover their critical strengths and limitations. Our benchmark study highlights that different model architectures and interpretability methods are better suited to specific scenarios. Negative samples derived from naturally inactive regions outperform synthetic sequences, whereas single-cell tasks require specialized models. Additionally, we demonstrate that interpretable sequence-function models can complement traditional sequence alignment methods in studying cross-species enhancer regulatory logic. We also provide a pipeline to help researchers select the optimal sequence-function prediction and interpretability algorithms.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 97%
- Cross-Species Prediction of Histone Modifications in Plants via Deep Learning 96%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 96%
Similar papers in this journal
- Assessing deep learning algorithms in cis-regulatory motif finding based on genomic sequencing data 97%
- FIRM: Flexible Integration of single-cell RNA-sequencing data for large-scale Multi-tissue cell atlas datasets 96%
- Cofea: correlation-based feature selection for single-cell chromatin accessibility data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.