HAETAE: A highly accurate and efficient epigenome transformer for tissue-specific histone modification prediction
Park, S.-J.; Im, S.-H.; Kim, S.-Y.; Kim, J.-Y.
Show abstract
While genomic models trained on four bases often fail to capture cell-type specificity, we introduce HAETAE, which integrates 5-methylcytosine from long-read sequencing into a 5-base framework. By explicitly modeling epigenetic context, HAETAE achieves state-of-the-art accuracy (>0.95) with orders of magnitude fewer parameters, challenging the prevailing scaling-law paradigm. Furthermore, HAETAE deciphers tissue-specific regulatory logic, as demonstrated by revealing the distinct, context-dependent functional impact of the TERT promoter mutation across diverse tissues.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Gapped-kmer sequence modeling robustly identifies regulatory vocabularies and distal enhancers conserved between evolutionarily distant mammals 96%
- Fine-mapping of nuclear compartments using ultra-deep Hi-C shows that active promoter and enhancer elements localize in the active A compartment even when adjacent sequences do not 95%
- Boosting the detection of enhancer-promoter loops via novel normalization methods for chromatin interaction data 95%
Similar papers in this journal
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 97%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 96%
- CREaTor: zero-shot cis-regulatory pattern modeling with attention mechanisms 96%
Similar papers in this journal
Similar papers in this journal
- Quantitative single cell 5hmC sequencing reveals non-canonical gene regulation by non-CG hydroxymethylation 96%
- Multi-omics integration and regulatory inference for unpaired single-cell data with a graph-linked unified embedding framework 95%
- Deaminase-assisted single-molecule and single-cell chromatin fiber sequencing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.