A novel deep learning method using contrastive learning enables interpretable outputs from relationships between gene expression and histone modification
Munif, A.; Datta, A.; Li, Z.; Ward, M.
Show abstract
Current methods for predicting gene expression from histone modifications rely on arbitrary binary classification thresholds such as the median to distinguish between high and low expression for genes. This approach lacks biological justification, creates dataset-dependent classifications, and ignores the relative regulatory relationships between genes that are often more biologically meaningful than absolute cutoffs. We introduce a novel pairwise ranking approach that compares relative expression levels between gene pairs based on their histone modification patterns, eliminating arbitrary threshold selection. Evaluation using the REMC E066 and GSE76344 datasets showed that our pairwise classification framework demonstrated consistent ranking performance in both datasets. The ablation study found that only a subset of histone modification types is necessary, giving evidence of redundancy in the biological code.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Molecular Group and Correlation Guided Structural Learning for Multi-Phenotype Prediction 95%
- SPCS: A Spatial and Pattern Combined Smoothing Method of Spatial Transcriptomic Expression 94%
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 94%
Similar papers in this journal
- Highly Effective Batch Effect Correction Method for RNA-seq Count Data 94%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 94%
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.