Feature Selection Using Lasso Regression Enhances Deep Learning Model Performance For Diagnosis Of Lung Cancer from Transcriptomic Data
GUHA, S.
Show abstract
Cancer is a genetic disease where gene mutations are pivotal in disease initiation and pathophysiology. The gene expression profile follows a specific pattern exclusive to each cancer which can be utilized for early and accurate diagnosis. Microarray techniques have emerged as powerful tools capable of simultaneously capturing the expression profiles of thousands of genes. However, because of the high dimensionality of the produced transcriptome data, analysis of the resulting datasets is challenging. Recent advancements in Artificial Intelligence (AI) techniques like Machine Learning (ML) and Deep Learning can be instrumental in efficiently processing these high-dimensional datasets. LASSO-regression is a ML technique that can help to rank the features which could help in feature selection leading to dimensionality reduction. Deep Learning is one of the most sophisticated ML techniques that can process high-dimensional data owing to the presence of more number of hidden layers in its neural network. We designed a Deep Neural Network (DNN) classifier model fused with a LASSO-based significant feature extractor for classifying the gene expression dataset containing a total of 51 samples of which 24 samples are of lung cancer patients and the remaining 27 samples are of normal individuals. A LASSO regression model was implemented to identify the genes that played a significant role in the classification. These significant gene expressions were then fed into a convergent Deep Neural Architecture. The classifier was trained with 70% data and the rest 30% was used for validation. The proposed classifier proved to provide better classification as compared to LASSO regression and DNN used individually. The two classes were classified with an average accuracy of 96.25%, average precision of 99.67%, average specificity of 99.45% and average sensitivity of 91.73% measured over thirty independent assessments. In some cases, the model was able to obtain a classification accuracy of 100%. This could open the path to early and better diagnosis of cancers from transcriptome data.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 96%
- Cardiac disease diagnosis based on GAN in case of missing data 96%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 96%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 97%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 97%
- Comparing protein-protein interaction networks of SARS-CoV-2 and (H1N1) influenza using topological features 95%
Similar papers in this journal
Similar papers in this journal
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 96%
- Enrichment analysis on regulatory subspaces: a novel direction for the superior description of cellular responses to SARS-CoV-2 96%
- A machine-learning Approach for Stress Detection Using Wearable Sensors in Free-living Environments 96%
Similar papers in this journal
- EnGRNT: Inference of gene regulatory networks using ensemble methods and topological feature extraction 96%
- Extensive In Silico Analysis of the Functional and Structural Consequences of SNPs in Human ARX Gene 94%
- Predicting the Epidemic Curve of the Coronavirus (SARS-CoV-2) Disease (COVID-19) Using Artificial Intelligence 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.