Classical statistical methods are powerful for the identification of novel targets for the survival of breast cancer patients
Insawang, B.; Ward, M.; Li, Z.; Datta, A.
Show abstract
Breast cancer is a leading cause of cancer-related deaths among women. The identification of survival-related target genes is critical for improving the prognosis and outcomes of breast cancer patients. Many methods have been applied to this investigation, such as bioinformatics and machine learning approaches, yet few targets identified from these approaches have been applied in clinics. Here, we present a novel approach by using classical statistical methods of Kolmogorov-Smirnov (KS) test and Jensen-Shannon (JS) divergence to analyse the survival time and gene expression data of breast cancer patients (BRCA) from The Cancer Genome Atlas (TCGA). These methods help compare the survival time distributions and differentiate patients into high and low-risk groups based on gene expression profiles. 1,124 survival-related genes were identified based on the KS test and 18 from JS divergence values. We also identified the optimal thresholds of the expression level of these target genes, which enabled the best separation of survival groups for all breast cancer patients and each subtype of breast cancer patients. These targets were further validated through bootstrapping to ensure that significant results are not due to chance. By comparing those survival targets from previous studies, we found two were novel targets, and two were consistent with previous reports. Overall, our study provides a novel approach for identifying survival targets for breast cancer patients by integrating a series of classical statistical methods, such as the KS test, JS divergence, and bootstrapping. Our approach could also be applied to identifying the survival targets for other cancer types and provide valuable insights into cancer research and clinical applications.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Directed Bayesian Networks established functional differences between breast cancer subtypes 95%
- The Impact of Variance in Carnitine Palmitoyltransferase-1 Expression on Breast Cancer Prognosis is Stratified by Clinical and Anthropometric Factors 94%
- Factors associated with breast lesions among women attending select teaching and referral health facilities in Kenya: A cross-sectional study 94%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 96%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 94%
- Automated and Manual Quantification of Tumour Cellularity in Digital Slides for Tumour Burden Assessment 92%
Similar papers in this journal
- Germline Mutation Analysis in Sporadic Breast Cancer Cases with Clinical Correlations 95%
- Computing Skin Cutaneous Melanoma Outcome from the HLA-alleles and Clinical Characteristics 92%
- Identification of Platform-Independent Diagnostic Biomarker Panel for Hepatocellular Carcinoma using Large-scale Transcriptomics Data 92%
Similar papers in this journal
- The role of KPNA2 mutations in breast cancer prognosis: A survey of publicly available databases 96%
- Effectiveness of educational intervention on breast cancer knowledge and breast self-examination among female university students in Bangladesh: a pre-post quasi-experimental one group study 94%
- PDAC-ANN: an artificial neural network to predict Pancreatic Ductal Adenocarcinoma based on gene expression 92%
Similar papers in this journal
- Immunohistochemical Profiling of Histone Modification Biomarkers Identifies Subtype-Specific Epigenetic Signatures and Potential Drug Targets in Breast Cancer 93%
- Multi-run Concrete Autoencoder to Identify Prognostic lncRNAs for 12 Cancers 93%
- Unveiling epigenetic regulatory elements associated with breast cancer development 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.