Back

OpEnHiMR: Optimization based Ensemble Model for Prediction of Histone Modification in Rice

Sinha, D.; Murmu, S.; Sarkar, A.; Yeasin, M.; Bhattacharjee, S.; Chaurasia, H.; Mishra, D. C.; Budhlakoti, N.; Srivastava, S.; Archak, S.; Kumari, A.; Jha, G. K.

2025-11-11 bioinformatics
10.1101/2025.11.10.687509 bioRxiv
Show abstract

Histone modifications are central to gene regulation, yet their systematic identification in plants remains limited due to the complexity of epigenomic landscapes. We present OpEnHiMR, an optimization-based ensemble learning framework for multiclass prediction of three key histone modifications, H3K4me3, H3K27me3, and H3K9ac, in rice. The framework integrates Support Vector Machines, Random Forest, and Gradient Boosting models, optimized via Ant Colony Optimization to maximize performance. Biologically meaningful features, including mononucleotide binary encoding, nucleotide chemical properties, GC content, and k-mer frequencies, were used for training after rigorous data curation and redundancy removal. OpEnHiMR achieved a classification accuracy of 77.54%, outperforming individual models and ensuring improved recall, specificity, and Matthews correlation coefficient. Model interpretability was enhanced using SHAP analysis, which highlighted critical sequence features influencing prediction outcomes. To promote community-wide adoption, a user-friendly webserver (https://dipro-sinha.shinyapps.io/OpEnHiMR/) and R package (https://cran.r-project.org/web/packages/OpEnHiMR/index.html) were developed. OpEnHiMR thus provides a scalable, accurate, and interpretable tool for histone modification prediction in plants, advancing epigenomics research and supporting data-driven crop improvement strategies.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.