Back

Modeling Joint Reference Regions for Omics Biomarkers in UK Biobank Proteomics

Pusparum, M.; Thas, O.; Ertaylan, G.

2026-09-04 health informatics
10.64898/2026.09.01.26361504 medRxiv
Show abstract

Conventional univariate reference intervals (UniRIs) are widely used to identify abnormal biomarker values, but they evaluate each biomarker independently and do not account for coordinated deviations between biomarkers. We developed and evaluated a joint reference region (JRR) framework for plasma proteomics data using the Olink proteomics dataset generated by the UK Biobank Pharma Proteomics Project, covering approximately 3,000 plasma proteins. JRRs were estimated for selected protein pairs in a healthy reference subset, while UniRIs were estimated separately for individual proteins using the nonparametric method. Both approaches were then evaluated in ICD-defined disease subsets. Biomarker discovery revealed sparse and heterogeneous disease--protein associations, with some proteins recurring across multiple phenotypes and others showing more disease-specific patterns. The added value of JRRs varied across diseases and protein pairs. Across evaluated protein pairs, 56.5\% showed higher sensitivity under the JRR framework than the UniRI of the first protein, and 47.3\% showed higher sensitivity than the UniRI of the second protein. At the disease level, the median proportion of protein pairs with improved JRR sensitivity was 0.57. JRRs were most informative when univariate detection was limited but a subset of diseased observations was flagged only by the joint region. These findings suggest that JRRs provide a complementary approach to UniRIs by capturing abnormal joint biomarker configurations in high-dimensional proteomics data.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.