Back

Gene-centered representation of coding and regulatory variation enables outcome prediction

Sokolova, K.; Kristensen, V.; Park, C. Y.; Theesfeld, C.; Mariani, L.; Kretzler, M.; Troyanskaya, O. G.; Cure Glomerulonephropathy (CureGN) Study Consortium,

2026-01-29 genomics
10.64898/2026.01.27.701808 bioRxiv
Show abstract

Integrating coding and regulatory variation into unified, interpretable representations remains a challenge in functional genomics. Current approaches either focus on common variants or analyze individual variants in isolation, missing the cumulative, cell-type-specific impact of both coding and noncoding variants on each gene. We present Volaria, a computational framework that integrates coding and regulatory genetic variation into unified, gene-centered representations for disease outcome prediction from whole-genome sequencing. Volaria leverages deep learning models to capture variant effects on cell-type-specific gene expression and integrates them with AI-predicted exonic variant pathogenicity to produce representations that capture the cumulative effect of genome-wide rare and common variation. Applied to whole genomes of individuals with rare glomerular diseases, Volaria predicts individual outcomes directly from germline sequence, demonstrating that structured, cell-type-aware representations capture predictive signals beyond population-based polygenic risk scores and unstructured representations. Importantly, the framework identifies context-specific biological mechanisms, providing interpretability that can be aligned with clinical measurements. By encoding genome-wide variation into compact and biologically grounded representations, Volaria provides a scalable foundation for genome interpretation and individualized outcome modeling from germline sequence, complementing phenotypic and clinical information in the future integrative frameworks.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.