A self-supervised DNA foundation model with collapse-resistant multimodal fusion
Chen, Y.
Show abstract
Genomic foundation models pretrained on DNA sequence have achieved strong performance across a range of tasks, but sequence-only representations cannot fully capture regulatory information reflected by additional DNA-centric modalities. Existing multimodal genomic models are often optimized for specific prediction tasks rather than for learning reusable embeddings shared across downstream analyses. However, directly fusing heterogeneous genomic modalities is challenging because sparse, peak-shaped regulatory signals and dense sequence representations have markedly different statistical structures, making naive multimodal alignment prone to degenerate near-zero solutions. We present a self-supervised DNA-centric multimodal foundation model that addresses this gap, integrating DNA sequence embeddings with local and global chromatin accessibility in a shared multimodal encoder to produce reusable window-level embeddings that support both masked reconstruction during pre-training and downstream prediction tasks. We diagnose this heterogeneous-modality alignment failure and show that global normalization substantially alleviates collapse, enabling effective joint learning across modalities. The resulting embeddings improve multiple downstream evaluations of regulatory function, including regulatory activity prediction, regulatory signal ranking and chromatin accessibility peak detection, achieving a 4.6-fold AUPRC improvement over the DNA-only baseline in peak detection, and further improving external validation on ClinVar, GTEx eQTL and PBMC caQTL datasets. The framework is extensible to additional regulatory modalities, providing a methodological basis for multimodal DNA foundation models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Prediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC 96%
- RegFormer: A Single-Cell Foundation Model Powered by Gene Regulatory Hierarchies 96%
- Empirical Bayes spline model learns multi-way genomic interactions from single cell 3D genome data 96%
Similar papers in this journal
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 96%
- Predicting enhancer-gene links from single-cell multi-omics data by integrating prior Hi-C information 95%
- Massively parallel reporter assay-informed modeling improves prediction of context-specific enhancer-gene regulatory interactions 94%
Similar papers in this journal
- Multi-scale deep tensor factorization learns a latent representation of the human epigenome 96%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 96%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.