Prioritizing Complex Disease Genes from Heterogeneous Public Databases
Gong, E. L.; Chen, J. Y.
Show abstract
BackgroundComplex human diseases are defined not only by sophisticated patterns of genetic variants/mutations upstream but also by many interplaying genes, RNAs, and proteins downstream. Analyzing multiple genomic and functional genomic data types to determine a short list of genes or molecules of interest is a common task called "gene prioritization" in biology. There are many statistical, biological, and bioinformatic methods developed to perform gene prioritization tasks. However, little research has been conducted to examine the relationships among the technique used, merged/separate use of each data modality, the gene lists network/pathway context, and various gene ranking/expansions. MethodsWe introduce a new analytical framework called "Gene Ranking and Iterative Prioritization based on Pathways" (GRIPP) to prioritize genes derived from different modalities. Multiple data sources, such as CBioPortal, PAGER, and COSMIC were used to compile the initial gene list. We used the PAGER software to expand the gene list based on biological pathways and the BEERE software to construct protein-protein interaction networks that include the gene list to rank order genes. We produced a final gene list for each data modality iteratively from an initial draft gene list, using glioblastoma multiform (GBM) as a case study. ConclusionWe demonstrated that GBM gene lists obtained from three modalities (differential gene expressions, gene mutations, and copy number alterations) and several data sources could be iteratively expanded and ranked using GRIPP. While integrating various modalities of data can be useful to generate an integrated ranked gene list related to any specific disease, the integration may also decrease the overall significance of ranked genes derived from specific data modalities. Therefore, we recommend carefully sorting and integrating gene lists according to each modality, such as gene mutations, epigenetic controls, or differential expressions, to procure modality-specific biological insights into the prioritized genes.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MonaGO: a novel Gene Ontology enrichment analysis visualisation system 95%
- pyCancerSig: subclassifying human cancer with comprehensive single nucleotide, structural and microsatellite mutational signature deconstruction from whole genome sequencing 94%
- HIHISIV: a database of gene expression in HIV and SIV host immune response 94%
Similar papers in this journal
Similar papers in this journal
- Analysis of Mutations in Precision Oncology using The Automated, Accurate, and User-Friendly Web Tool PredictONCO 95%
- iMDA-BN: Identification of miRNA-Disease Associations based on the Biological Network and Graph Embedding Algorithm 94%
- eDAVE - extension of GDC Data Analysis, Visualization, and Exploration Tools 94%
Similar papers in this journal
Similar papers in this journal
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 95%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 95%
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.