E.PAGE: A curated database and enrichment tool to predict modules associated with gene-environment interactions
Muralidharan, S.; Zahir, F.; Mehdi, A. M.
Show abstract
BackgroundThe purpose of this study was to manually and semi-automatically curate a database and develop an R package that will provide a comprehensive resource to uncover associations between biological processes and environmental factors in health and disease. We followed a two-step process to achieve the objectives of this study. First, we conducted a systematic review of existing gene expression datasets to identify those with integrated genomic and environmental factors. This enabled us to curate a comprehensive genomic-environmental database for four key environmental factors (smoking, diet, infections and toxic chemicals) associated with various autoimmune and chronic conditions. Second, we developed a statistical analysis package that allows users to interrogate the relationships between differentially expressed genes and environmental factors under different disease conditions. ResultsThe initial database search run on the Gene Expression Omnibus (GEO) and the Molecular Signature Database (MSigDB) retrieved a total of 90,018 articles. After title and abstract screening against pre-set criteria, a total of 186 studies were selected. From those, 243 individual sets of genes, or "gene modules", were obtained. We then curated a database containing four environmental factors, namely cigarette smoking, diet, infections and toxic chemicals, along with a total of 25789 genes that had an association with one or more of these gene modules. In six case studies, the database and statistical analysis package were then tested with lists of differentially expressed genes obtained from the published literature related to type 1 diabetes, rheumatoid arthritis, small cell lung cancer, COVID-19, cobalt exposure and smoking. On testing, we uncovered statistically enriched biological processes, which revealed pathways associated with environmental factors and the genes. ConclusionsA novel curated database and software tool is provided as an R Package. Users can enter a list of genes to discover associated environmental factors under various disease conditions.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Systems biology analysis of publicly available transcriptomic data reveals a critical link between AKR1B10 gene expression, smoking and occurrence of lung cancer 93%
- Association of indoor use of pesticides with CKD of unknown origin 93%
- Transcriptomic profiling of PBDE-exposed HepaRG cells unveils critical lncRNA- PCG pairs involved in intermediary metabolism 93%
Similar papers in this journal
- Structural variability, expression profile and pharmacogenetics properties of TMPRSS2 gene as a potential target for COVID-19 therapy 92%
- Unveiling sex-based differences in the effects of alcohol abuse: a comprehensive functional meta-analysis of transcriptomic studies 92%
- Integrated Analysis of Tissue-specific Gene Expression in Diabetes by Tensor Decomposition Can Identify Possible Associated Diseases. 91%
Similar papers in this journal
- Genetic regulators of mineral amount in Nelore cattle muscle predicted by a new co-expression and regulatory impact factor approach 93%
- Effects of repetitive Iodine Thyroid Blocking on the Development of the Foetal Brain and Thyroid in rats: a Systems Biology approach 92%
- Comparative transcriptome analyses reveal genes associated with SARS-CoV-2 infection of human lung epithelial cells 92%
Similar papers in this journal
- Developmental pyrethroid exposure disrupts molecular pathways for MAP kinase and circadian rhythms in mouse brain 89%
- Transcriptional Dynamics of Sleep Deprivation and Subsequent Recovery Sleep in the Male Mouse Cortex 88%
- Maternal-fetal interfaces transcriptome changes associated with placental insufficiency and a novel gene therapy intervention 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.