Explainable-AI to Discover Associated Genes for Classifying Hepato-cellular Carcinoma from High-dimensional Data
Hasan, E.; Mostafa, F.; Hossain, M. S.; Loftin, J.
Show abstract
Knowledge-based interpretations are essential for understanding the omic data set because of its nature, such as high dimension and hidden biological information in genes. When analyzing gene expression data with many genes and few samples, the main problem is to separate disease-related information from a vast quantity of redundant data and noise. This paper uses a reliable framework to determine important genes for discovering Hepato-cellular Carcinoma (HCC) from micro-array analysis and eliminating redundant and unnecessary genes through gene selection. Several machine learning models were applied to find significant predictors responsible for HCC. As we know, classification algorithms such as Random Forest, Naive Bayes classifier, or a k-Nearest Neighbor classifier can help us to classify HCC from responsible genes. Random Forests shows 96.53% accuracy with p < 0.00001, which is better than other discussed Machine Learning(ML) approaches. Each gene is not responsible for a particular patient. Since ML approaches are like black boxes and people/practitioners do not rely on them sometimes. Artificial Intelligence(AI) technologies with high optimization interoperability shed light on what is happening inside these systems and aid in the detection of potential problems; including causality, information leakage, model bias, and robustness when determining responsible genes for a specific patient with a high probability score of almost 0.99, from one of the samples mentioned in this study.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Stability of feature selection utilizing Graph Convolutional Neural Network and Layer-wise Relevance Propagation 95%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 94%
- Fenchel duality of Cox partial likelihood and its application in survival kernel learning 94%
Similar papers in this journal
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 97%
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 96%
- Unsupervised Discovery of Risk Profiles on Negative and Positive COVID-19 Hospitalized Patients 96%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 96%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 96%
- Comparing protein-protein interaction networks of SARS-CoV-2 and (H1N1) influenza using topological features 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.