Systematic Normalization with Multiple Housekeeping Genes for the Discovery of Genetic Dependencies in Cancer
Bonham-Carter, O.; Thu, Y. M.
Show abstract
Cancer results from complex interactions between genes that are misregulated. Although our understanding of the contribution of single genes to cancer is expansive, the interplay between genes in the context of this devastating disease remains to be understood. Using the Genomic Data Commons Data Portal through National Cancer Institute, we randomly selected ten data sets of breast cancer gene expression, acquired by RNA sequencing to be subjected to a computational method for the exploration of genetic interactions at a large scale. We focused on genes that suppress genome instability (GIS genes) since function or expression of these genes is often altered in cancer. In this paper, we show how to discover pairs of genes whose expressions demonstrate patterns of correlation. To ensure an inter-comparison across data sets, we tested statistical normalization approaches derived from the expression of randomly selected single housekeeping genes, or from the average of three. In addition, we systematically selected ten housekeeping genes for the purpose of normalization. Using normalized expression data, we determined R2 values from linear models for all possible pairs of GIS genes and presented our results using heatmaps. Despite the heterogeneity of data, we observed that multiple gene normalization revealed more consistent correlations between pairs of genes, compared to using single gene expressions. We also noted that multiple gene normalization using ten genes outperformed normalization using three randomly selected genes. Since this study uses gene expression data from cancer tissues and begins to address the reproducibility of correlation between two genes, it complements other efforts to identify gene pairs that co-express in cancer cell lines. In the future, we plan to define consistent genetic correlations by using gene expression data derived from different types of cancer and multiple gene normalization. CCS CONCEPTSO_LIApplied computing [->] Computational biology. C_LI ACM Reference FormatOliver Bonham-Carter and Yee Mon Thu. 2019. Systematic Normalization with Multiple Housekeeping Genes for the Discovery of Genetic Dependencies in Cancer. In Niagara Falls, New York. ACM, New York, NY, USA, 10 pages. https://doi.org/10.1145/nnnnnnn.nnnnnnn
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 96%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 95%
- Identification of miRNA signatures for kidney renal clear cell carcinoma using the tensor-decomposition method 95%
Similar papers in this journal
- Normalization of RNA-Seq Data using Adaptive Trimmed Mean with Multi-reference 96%
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 96%
- Feature Extraction Approaches for Biological Sequences: A Comparative Study of Mathematical Models 95%
Similar papers in this journal
Similar papers in this journal
- The Aliment to Bodily Condition knowledgebase (ABCkb): A database connecting plants and human health 92%
- Effect of Heat inactivation and bulk lysis on Real-Time Reverse Transcription PCR Detection of the SARS-COV-2: An Experimental Study 92%
- First report of reference guided genome assembly of Black Bengal goat (Capra hircus) 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.