Back

Explainable AI identifies H3K18ac as a new marker of active enhancers

Maqsood, K.; Polvora Brandao, D.; Wolfe, J.; Kayyar, B.; Clinciu, C. G.; Bosnea, R. A.; Grant, O. A.; Boulet, F.; Bell, C. G.; Ficz, G.; Madapura, P.; Hagras, H.; Zabet, N. R.

2026-06-11 genomics
10.64898/2026.06.09.731088 bioRxiv
Show abstract

Enhancers are non-coding regions of DNA that regulate gene transcription, yet the mechanisms underlying enhancer activity remain incompletely understood. Despite extensive experimental and computational efforts, we still lack accurate enhancer maps in many human cells, tissues and disease contexts. Here, we developed several Artificial Intelligence (AI) models (Convolutional Neural Networks (CNN), XGBoost, Logistic Regression (LR) and an eXplainable Artificial Intelligence type2 Fuzzy Logic based System (type2-FLS)) to predict enhancers across different human and mouse cell lines. While all models display high accuracy in the cell lines they were trained on, our results confirmed that type2-FLS, and, partially, CNN, LR and XGBoost perform consistently well in cell lines unseen during training, supporting the generalisation of the models. Most importantly, type2-FLS identified H3K18ac as an important enhancer mark along with many novel putative enhancers, which display the same epigenetic signatures as experimentally identified ones. We have validated some of these novel enhancers by both global epigenetic perturbations and directed enhancer epigenetic rewriting (CRISPRi). Interestingly, seven epigenetic marks in humans and five in mouse are sufficient to annotate enhancers without losing accuracy. Overall, we have deciphered the epigenetic code of mammalian enhancers and annotated enhancers in multiple human and mouse cell lines.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Genome Biology
637 papers in training set
Top 0.2%
22.0%
2
Nucleic Acids Research
1281 papers in training set
Top 2%
11.1%
3
Nature Communications
5641 papers in training set
Top 18%
9.7%
4
Genome Research
468 papers in training set
Top 0.5%
7.9%
50% of probability mass above
5
Epigenetics & Chromatin
10 papers in training set
Top 0.1%
5.6%
6
BMC Genomics
406 papers in training set
Top 1%
4.9%
7
Scientific Reports
3612 papers in training set
Top 33%
3.2%
8
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.2%
9
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.7%
10
eLife
5828 papers in training set
Top 39%
2.7%
11
Bioinformatics
1204 papers in training set
Top 6%
2.1%
12
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.1%
1.7%
13
PLOS Computational Biology
1863 papers in training set
Top 17%
1.1%
14
Cell Reports
1498 papers in training set
Top 23%
1.1%
15
Epigenetics & Chromatin
42 papers in training set
Top 0.4%
1.1%
16
Nature Genetics
286 papers in training set
Top 4%
1.1%
17
Nature Structural & Molecular Biology
18 papers in training set
Top 0.3%
1.1%
18
Nature
645 papers in training set
Top 9%
1.0%
19
PLOS ONE
5266 papers in training set
Top 60%
0.9%
20
Cell Reports Methods
165 papers in training set
Top 4%
0.8%
21
Epigenetics
50 papers in training set
Top 0.6%
0.8%
22
Genome Medicine
183 papers in training set
Top 5%
0.8%
23
Molecular Systems Biology
162 papers in training set
Top 4%
0.6%
24
Molecular Cell
350 papers in training set
Top 5%
0.6%
25
Cell Genomics
172 papers in training set
Top 4%
0.6%
26
Mobile DNA
31 papers in training set
Top 0.3%
0.6%
27
Methods
34 papers in training set
Top 0.8%
0.6%
28
Nature Methods
385 papers in training set
Top 7%
0.6%
29
BMC Bioinformatics
457 papers in training set
Top 6%
0.6%