Robust Hierarchical Co-clustering to Explore Toxicogenomic Biomarkers and Their Regulatory Doses of Chemical Compounds
Hasan, M. N.; Badsha, M. B.; Mollah, M. N. H.
Show abstract
Toxicogenomics combines high throughput molecular technologies with statistical and machine learning approaches to discover a similar group of doses of chemical compounds (DCCs) and genes to explore toxicogenomic biomarkers and their regulatory DCCs. This is also very important in the toxicity study of environmental stressors, synthetic chemicals and drug discovery and development process. Different clustering algorithms are concerned with the discovering of interesting clusters/groups of row or column entities of a dataset. Among those hierarchical clustering (HC) and logistic probabilistic hidden variable model (LPHVM) can identify toxicogenomic biomarkers and their regulatory DCCs forming co-cluster. However, the HC method is very sensitive to outlying observations. On the other hand, though LPHVM is a robust approach, it consumes more time for calculation since it is Expectation-Maximization (EM) based iterative approach. Additionally, the LPHVM creates artificiality problem taking absolute value of the data matrix. Therefore, to overcome these problems in this paper, we proposed a robust hierarchical co-clustering (RHCOC) algorithm to co-cluster genes and DCCs simultaneously with a view to explore toxicogenomic biomarkers and their regulatory DCCs. The performance of the proposed RHCOC algorithm over the conventional HC for clustering genes and DCCs of toxicogenomic data has been investigated based on the simulation study. The results of the simulation study have shown that the RHCOC approaches produce far lower clustering error rate (ER) than the conventional HC approaches in presence of outlying observations in the dataset. Otherwise they perform equally in absence of outlier in the dataset. To explore biomarker co-clusters consisting of toxicogenomic biomarker genes and their regulatory DCCs we used control chart for individual measurement (CCIM). We have also investigated the performance of the proposed approach in the case of the pathway level real life fold change gene expression (FCGE) toxicogenomic data analysis. The biomarker co-clusters consisting of toxicogenomic biomarker genes and their regulatory DCCs and biomarker genes explored by the proposed approaches have been validated by the literature and functional annotation. Our method is implemented in R package "rhcoclust" available on github (https://github.com/mdbahadur/rhcoclust).
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 94%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 94%
- Unsupervised tensor decomposition-based method to extract candidate transcription factors as histone modification bookmarks in post-mitotic transcriptional reactivation 93%
Similar papers in this journal
Similar papers in this journal
- Combining Multi-Dimensional Molecular Fingerprints to Predict hERG Cardiotoxicity of Compounds 94%
- Enrichment analysis on regulatory subspaces: a novel direction for the superior description of cellular responses to SARS-CoV-2 93%
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.