Back

Stratification of Risk of Progression to Colectomy in Ulcerative Colitis using Measured and Predicted Gene Expression

Mo, A.; Nagpal, S.; Gettler, K.; Haritunians, T.; Giri, M.; Haberman, Y.; Karns, R.; Prince, J.; Arafat, D.; Hsu, N.-Y.; Chuang, L.-S.; Argmann, C.; Kasarskis, A.; Suarez-Farinas, M.; Gotman, N.; Mengesha, E.; Venkateswaran, S.; Rufo, P. A.; Baker, S. S.; Sauer, C. G.; Markowitz, J.; Pfefferkorn, M. D.; Rosh, J. R.; Boyle, B. M.; Mack, D. R.; Baldassano, R. N.; Shah, S.; LeLeiko, N. S.; Heyman, M. B.; Griffiths, A. M.; Patel, A. S.; Noe, J. D.; Thomas, S. D.; Aronow, B. J.; Walters, T. D.; McGovern, D. P.; Hyams, J. S.; Kugathasan, S.; Cho, J.; Denson, L. A.; Gibson, G.

2021-04-04 genetics
10.1101/2021.04.02.438187 bioRxiv
Show abstract

An important goal of clinical genomics is to be able to estimate the risk of adverse disease outcomes. Between 5% and 10% of ulcerative colitis (UC) patients require colectomy within five years of diagnosis, but polygenic risk scores (PRS) utilizing findings from GWAS are unable to provide meaningful prediction of this adverse status. By contrast, in Crohns disease, gene expression profiling of GWAS-significant genes does provide some stratification of risk of progression to complicated disease in the form of a Transcriptional Risk Score (TRS). Here we demonstrate that both measured (TRS) and polygenic predicted gene expression (PPTRS) identify UC patients at 5-fold elevated risk of colectomy with data from the PROTECT clinical trial and UK Biobank population cohort studies, independently replicated in an NIDDK-IBDGC dataset. Prediction of gene expression from relatively small transcriptome datasets can thus be used in conjunction with transcriptome-wide association studies to stratify risk of disease complications.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.