Back

An interpretable omnigenic neural network architecture for the human genome

Upmeier zu Belzen, J.; Arnoldt, L.; Hollmann, N.; Herrmann, L.; Nguyen, K. M.; Eckhoff, L.; Kohleick, L.; Abou Ghaloun, S.; Schmidt, H.; Hegselmann, S.; Theis, F. J.; Buergel, T.; Steinfeldt, J.; Wild, B.; Eils, R.

2026-07-30 genetic and genomic medicine
10.64898/2026.07.28.26359187 medRxiv
Show abstract

Genetic prediction of complex phenotypes typically relies on additive linear models, which scale well but cannot capture non-additive effects or deeply integrate molecular and clinical data. Domain-specific neural networks have driven advances in images, text, and other modalities, but genome-scale neural networks remain challenging because genotypes are sparse and high-dimensional, effective sample sizes are limited, and generic architectures lack interpretability. Here, we introduce the omnigenic neural network, a biologically structured architecture inspired by the omnigenic model of complex traits. The model learns hierarchical representations of biological processes, accommodates multimodal inputs, supports transfer learning, and enables multitask prediction. Models trained in the UK Biobank and evaluated in the All of Us cohort for ischemic heart disease, type 2 diabetes, and schizophrenia outperformed published PGS Catalog and PRS-CSx scores. A multitask model trained across 36 cardiovascular endpoints further outperformed corresponding single-phenotype models and baselines. The architecture provides systems-level interpretability by quantifying the contributions of biological processes, which were consistent with established disease mechanisms. It also captures non-linear interactions between variants. Analysis of these interactions using Integrated Hessians revealed patterns concordant with previously reported epistatic associations. Together, these findings establish the omnigenic neural network as a flexible framework for interpretable, multimodal, and multitask genomic prediction.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.