Back

OmicFormer: a statistical priors-informed transformer for accurate and generalizable omics prediction of diseases and complex traits

Jiang, H.; Yang, C.; Qin, M.; You, J.; Feng, J.; Yu, J.-T.; Cheng, W.; Gong, W.

2026-07-10 health informatics
10.64898/2026.07.06.26357359 medRxiv
Show abstract

Precision medicine faces a critical challenge in translating high-dimensional omics data into robust disease predictions across diverse populations. Current approaches often fail under distribution shifts, partly due to their inability to encode complex biological feature dependencies. We present OmicFormer, a Transformer-based architecture that embeds two complementary statistical priors, i.e., feature-label associations and feature-feature dependencies, directly into its representation learning. This design captures local and long-range omic interactions often missed by conventional methods. Analyzing 500,000 UK Biobank participants, OmicFormer significantly outperforms strong baselines across 450 disease and 900 trait prediction tasks , with substantial gains spanning diverse metabolic, neurological, cardiovascular, and gastrointestinal conditions, alongside enhanced prediction of circulating metabolites, bone density traits, and retinal imaging biomarkers. Crucially, OmicFormer demonstrates robust generalization, achieving a substantial improvement over tree-based methods in an independent proteomics cohort across 19 diseases (GNPC, N=7,289), and outperforming tree-based models across 50 multi-site neuroimaging sites (N=4,728) for autism and schizophrenia classification. By explicitly embedding statistical structure, OmicFormer provides an interpretable and generalizable foundation for omics-based precision medicine.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.