Comprehensive biological interpretation of gene signatures using semantic distributed representation
Okuzono, Y.; Hoshino, T.
Show abstract
Recent rise of microarray and next-generation sequencing in genome-related fields has simplified obtaining gene expression data at whole gene level, and biological interpretation of gene signatures related to life phenomena and diseases has become very important. However, the conventional method is numerical comparison of gene signature, pathway, and gene ontology (GO) overlap and distribution bias, and it is not possible to compare the specificity and importance of genes contained in gene signatures as humans do. This study proposes the gene signature vector (GsVec), a unique method for interpreting gene signatures that clarifies the semantic relationship between gene signatures by incorporating a method of distributed document representation from natural language processing (NLP). In proposed algorithm, a gene-topic vector is created by multiplying the feature vector based on the genes distributed representation by the probability of the gene signature topic and the low frequency of occurrence of the corresponding gene in all gene signatures. These vectors are concatenated for genes included in each gene signature to create a signature vector. The degrees of similarity between signature vectors are obtained from the cosine distances, and the levels of relevance between gene signatures are quantified. Using the above algorithm, GsVec learned approximately 5,000 types of canonical pathway and GO biological process gene signatures published in the Molecular Signatures Database (MSigDB). Then, validation of the pathway database BioCarta with known biological significance and validation using actual gene expression data (differentially expressed genes) were performed, and both were able to obtain biologically valid results. In addition, the results compared with the pathway enrichment analysis in Fishers exact test used in the conventional method resulted in equivalent or more biologically valid signatures. Furthermore, although NLP is generally developed in Python, GsVec can execute the entire process in only the R language, the main language of bioinformatics.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 96%
- Differentially Expressed Heterogeneous Overdispersion Genes Testing for Count Data 95%
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 95%
Similar papers in this journal
- SPCS: A Spatial and Pattern Combined Smoothing Method of Spatial Transcriptomic Expression 97%
- An in silico approach to identification, categorization and prediction of nucleic acid binding proteins 96%
- MeSHHeading2vec: A new method for representing MeSH headings as feature vectors based on graph embedding algorithm 96%
Similar papers in this journal
- Enrichment analysis on regulatory subspaces: a novel direction for the superior description of cellular responses to SARS-CoV-2 97%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 96%
- Unsupervised Discovery of Risk Profiles on Negative and Positive COVID-19 Hospitalized Patients 95%
Similar papers in this journal
- iMDA-BN: Identification of miRNA-Disease Associations based on the Biological Network and Graph Embedding Algorithm 98%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 95%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.