Back

Challenges in predicting chromatin accessibility differences between species

Stephen, A. Z. M.; Raje, A.; Sestili, H. H.; Wirthlin, M. E.; Lawler, A. J.; Brown, A. R.; Stauffer, W. R.; Pfenning, A. R.; Kaplow, I. M.

2025-11-10 genomics
10.1101/2025.11.09.687449 bioRxiv
Show abstract

Enhancers are transcriptional regulatory elements that help drive phenotypic diversity, yet they often undergo rapid sequence evolution despite functional conservation, posing a challenge for predicting their function across species. Machine learning models that predict quantitative enhancer activity using DNA sequence have not previously been evaluated for their ability to predict quantitative differences across orthologous regions. Here, we trained convolutional neural networks (CNNs) on a regression task to predict chromatin accessibility, which is a proxy for enhancer activity, in the liver across five mammals, and we developed a novel framework to evaluate cross-species performance. We demonstrated that training on multiple species improves model generalization to both species used in training and held-out species. However, the models consistently achieved poor performance in predicting quantitative differences in accessibility between species at orthologous regions. Our study highlights the challenges in using regression models to predict chromatin accessibility changes between species.

Published in NAR Genomics and Bioinformatics (predicted rank #2) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.