OpEnHiMR: Optimization based Ensemble Model for Prediction of Histone Modification in Rice
Sinha, D.; Murmu, S.; Sarkar, A.; Yeasin, M.; Bhattacharjee, S.; Chaurasia, H.; Mishra, D. C.; Budhlakoti, N.; Srivastava, S.; Archak, S.; Kumari, A.; Jha, G. K.
Show abstract
Histone modifications are central to gene regulation, yet their systematic identification in plants remains limited due to the complexity of epigenomic landscapes. We present OpEnHiMR, an optimization-based ensemble learning framework for multiclass prediction of three key histone modifications, H3K4me3, H3K27me3, and H3K9ac, in rice. The framework integrates Support Vector Machines, Random Forest, and Gradient Boosting models, optimized via Ant Colony Optimization to maximize performance. Biologically meaningful features, including mononucleotide binary encoding, nucleotide chemical properties, GC content, and k-mer frequencies, were used for training after rigorous data curation and redundancy removal. OpEnHiMR achieved a classification accuracy of 77.54%, outperforming individual models and ensuring improved recall, specificity, and Matthews correlation coefficient. Model interpretability was enhanced using SHAP analysis, which highlighted critical sequence features influencing prediction outcomes. To promote community-wide adoption, a user-friendly webserver (https://dipro-sinha.shinyapps.io/OpEnHiMR/) and R package (https://cran.r-project.org/web/packages/OpEnHiMR/index.html) were developed. OpEnHiMR thus provides a scalable, accurate, and interpretable tool for histone modification prediction in plants, advancing epigenomics research and supporting data-driven crop improvement strategies.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Probabilistic modeling methods for cell-free DNA methylation based cancer classification 95%
- Shrinkage estimation of gene interaction networks in single-cell RNA sequencing data 95%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 95%
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 96%
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 96%
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 96%
Similar papers in this journal
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 96%
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 95%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 94%
Similar papers in this journal
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 96%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 96%
- A cautionary tale about properly vetting datasets used in supervised learning predicting metabolic pathway involvement 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.