Back

Effects of underlying gene-regulation network structure on prediction accuracy in high-dimensional regression

Okinaga, Y.; Kyogoku, D.; Kondo, S.; Nagano, A. J.; Hirose, K.

2020-09-12 bioinformatics
10.1101/2020.09.11.293456 bioRxiv
Show abstract

MotivationThe least absolute shrinkage and selection operator (lasso) and principal component regression (PCR) are popular methods of estimating traits from high-dimensional omics data, such as transcriptomes. The prediction accuracy of these estimation methods is highly dependent on the covariance structure, which is characterized by gene regulation networks. However, the manner in which the structure of a gene regulation network together with the sample size affects prediction accuracy has not yet been sufficiently investigated. In this study, Monte Carlo simulations are conducted to investigate the prediction accuracy for several network structures under various sample sizes. ResultsWhen the gene regulation network was random graph, the simulation indicated that models with high estimation accuracy could be achieved with small sample sizes. However, a real gene regulation network is likely to exhibit a scale-free structure. In such cases, the simulation indicated that a relatively large number of observations is required to accurately predict traits from a transcriptome. Availability and implementationSource code at https://github.com/keihirose/simrnet Contacthirose@imi.kyushu-u.ac.jp

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.