Small Patient Datasets Reveal Genetic Drivers of Non-Small Cell Lung Cancer Subtypes using a Novel Machine Learning Approach
Cook, M. P.; Qorri, B.; Baskar, A.; Ziauddin, J.; Pani, L.; Bushan Yenkanchi, S.; Geraci, J.
Show abstract
BackgroundThere are many small datasets of significant value in the medical space that are being underutilized. Due to the heterogeneity of complex disorders found in oncology, systems capable of discovering patient subpopulations while elucidating etiologies is of great value as it can indicate leads for innovative drug discovery and development. Materials and MethodsHere, we report on a machine intelligence-based study that utilized a combination of two small non-small cell lung cancer (NSCLC) datasets consisting of 58 samples of adenocarcinoma (ADC) and squamous cell carcinoma (SCC) and 45 samples (GSE18842). Utilizing a set of standard machine learning (ML) methods which are described in this paper, we were able to uncover subpopulations of ADC and SCC while simultaneously extracting which genes, in combination, were significantly involved in defining the subpopulations. We also utilized a proprietary interactive hypothesis-generating method designed to work with machine learning methods, which provided us with an alternative way of pinpointing the most important combination of variables. The discovered gene expression variables were used to train ML models. This allowed us to create methods using standard methods and to also validate our in-house methods for heterogeneous patient populations, as is often found in oncology. ResultsUsing these methods, we were able to uncover genes implicated by other methods and accurately discover known subpopulations without being asked, such as different levels of aggressiveness within the SCC and ADC subtypes. Furthermore, PIGX was a novel gene implicated in this study that warrants further study due to its role in breast cancer proliferation. ConclusionHere we demonstrate the ability to learn from small datasets and reveal well-established properties of NSCLC. This demonstrates the utility for machine learning techniques to reveal potential genes of interest, even from small data sets, and thus the driving factors behind subpopulations of patients.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PanClassif: Improving pan cancer classification of single cell RNA-seq using machine learning 91%
- Minimum Error Calibration and Normalization for Genomic Copy Number Analysis 90%
- Gene expression profiles and pathway enrichment analysis to identification of differentially expressed gene and signaling pathways in epithelial ovarian cancer based on high-throughput RNA-seq data 90%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 94%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 94%
- Lung biopsy cells transcriptional landscape from COVID-19 patient stratified lung injury in SARS-CoV-2 infection through impaired pulmonary surfactant metabolism 92%
Similar papers in this journal
- High Expression of Glycolytic Genes in Clinical Glioblastoma Patients Correlates with Lower Survival 90%
- Renal carcinoma is associated with increased risk of coronavirus infections 89%
- Analytical validation and clinical utilization of K-4CARE TM : a comprehensive genomic profiling assay with personalized MRD detection 88%
Similar papers in this journal
- Cancer type classification in liquid biopsies based on sparse mutational profiles enabled through data augmentation and integration 91%
- Generating complex explanations for artificial intelligence models: an application to clinical data on severe mental illness 88%
- Investigation of cell mechanics and migration on DDR2-expressing neuroblastoma cell line 87%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.