LASSO Based Analysis for Prediction of Prognostic Signature Genes Associated with Breast Cancer
Guha, S.
Show abstract
BackgroundCancer is a genetic disease, where gene alterations play a significant role in the disease onset and pathogenesis. Analysis of the underlying gene interaction pathways could reveal new biomarkers and could also potentially help in the development of targeted drugs for therapeutics. Microarray techniques have emerged as powerful tools capable of simultaneously measuring the expression levels of thousands of genes, making them invaluable in cancer biology research. However, the processing of the resultant datasets poses significant challenges due to their high dimensionality. Also, feature extraction becomes essential to discern the crucial features within these extensive datasets. To mitigate these difficulties advanced computational techniques like Machine Learning (ML) could be instrumental. LASSO-regression-based classification is an advanced ML technique that can help in feature selection by evaluating individual parameters like genes. MethodsThis study focuses on uncovering key prognostic genes for breast cancer using a combination of LASSO regression-based classifier and statistical bioinformatics models. Differentially expressed genes (DEGs) were identified using the "Limma" package in R, and significant genes were further filtered using the LASSO-based classifier significance coefficient. Genes common to both methods were considered as the focus of this study. Additionally, Protein-Protein Interaction (PPI) networks of these key genes were constructed using STRING, and hub genes, significant modules, and associated genes were identified using Cytoscape. ResultsThis study identified CCR8, CXCL11, CCL23, CCL24, CCL28, and CCL21 as signature prognostic genes for breast cancer, revealing a strong association between chemokines and breast cancer pathogenesis. Extensive literature searches were conducted to validate and confirm their prognostic significance in the disease. ConclusionThese findings are pivotal for enhancing our comprehension of the pathways involved in breast cancer. Additionally, they hold promise as novel biomarkers for diagnostic purposes and may also reveal significant therapeutic targets for the management of breast cancer.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Directed Bayesian Networks established functional differences between breast cancer subtypes 96%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 95%
- Machine learning based prediction of recurrence after curative resection for rectal cancer 95%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 98%
- Discovering Key Transcriptomic Regulators in Pancreatic Ductal Adenocarcinoma using Dirichlet Process Gaussian Mixture Model 97%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 95%
Similar papers in this journal
Similar papers in this journal
- Loss of CHGA protein as a potential biomarker for colon cancer diagnosis: a study on biomarker discovery by machine learning and confirmation by immunohistochemistry in colorectal cancer tissue microarrays 96%
- A novel molecular analysis approach in colorectal cancer suggests new treatment opportunities 94%
- Hormone Receptor-status Prediction in Breast Cancer Using Gene Expression Profiles and Their Macroscopic Landscape 94%
Similar papers in this journal
- Systems biomedicine of primary and metastatic colorectal cancer reveals potential therapeutic targets 97%
- Patient stratification of clear cell renal cell carcinoma using the global transcription factor activity landscape derived from RNA-seq data 96%
- Absence of Glutathione S-Transferase Theta 1 Gene is Significantly Associated with Breast Cancer Susceptibility in Pakistani Population and Poor Overall Survival in Breast Cancer Patients: A Case-Control and Case Series Analysis 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.