Comparing statistical learning methods for complex trait prediction from gene expression
Klimkowski Arango, N.; Morgante, F.
Show abstract
Accurate prediction of complex traits is an important task in quantitative genetics that has become increasingly relevant for personalized medicine. Genotypes have traditionally been used for trait prediction using a variety of methods such as mixed models, Bayesian methods, penalized regressions, dimension reductions, and machine learning methods. Recent studies have shown that gene expression levels can produce higher prediction accuracy than genotypes. However, only a few prediction methods were used in these studies. Thus, a comprehensive assessment of methods is needed to fully evaluate the potential of gene expression as a predictor of complex trait phenotypes. Here, we used data from the Drosophila Genetic Reference Panel (DGRP) to compare the ability of several existing statistical learning methods to predict starvation resistance from gene expression in the two sexes separately. The methods considered differ in assumptions about the distribution of gene effect sizes - ranging from models that assume that every gene affects the trait to more sparse models - and their ability to capture gene-gene interactions. We also used functional annotation (i.e., Gene Ontology (GO)) as an external source of biological information to inform prediction models. The results show that differences in prediction accuracy between methods exist, although they are generally not large. Methods performing variable selection gave higher accuracy in females while methods assuming a more polygenic architecture performed better in males. Incorporating GO annotations further improved prediction accuracy for a few GO terms of biological significance. Biological significance extended to the genes underlying highly predictive GO terms with different genes emerging between sexes. Notably, the Insulin-like Receptor (InR) was prevalent across methods and sexes. Our results confirmed the potential of transcriptomic prediction and highlighted the importance of selecting appropriate methods and strategies in order to achieve accurate predictions.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genomic prediction informed by biological processes expands our understanding of the genetic architecture underlying free amino acid traits in dry Arabidopsis seeds 93%
- Leveraging multiple layers of data to predict Drosophila complex traits 93%
- Adding gene transcripts into genomic prediction improves accuracy and reveals sampling time dependence 92%
Similar papers in this journal
- Gene co-expression networks identify novel candidate genes for moulting and development in the Atlantic salmon louse (Lepeophtheirus salmonis) 94%
- A Nextflow pipeline for molecular quantitative trait loci mapping in small sample size datasets with an application in Atlantic salmon 94%
- Diverse biological processes coordinate the transcriptional response to nutritional changes in a Drosophila melanogaster multiparent population 94%
Similar papers in this journal
- CoVar: A generalizable machine learning approach to identify the coordinated regulators driving variational gene expression 94%
- Detection of genes with differential expression dispersion unravels the role of autophagy in cancer progression 94%
- Sparse Multitask group Lasso for Genome-Wide Association Studies 93%
Similar papers in this journal
- Developmental transcriptomic analysis of the cave-dwelling crustacean, Asellus aquaticus 92%
- Integrated Analysis of Tissue-specific Gene Expression in Diabetes by Tensor Decomposition Can Identify Possible Associated Diseases. 92%
- Draft Genomes of two Artocarpus plants, Jackfruit (A. heterophyllus) and Breadfruit (A. altilis) 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.