Back

Quantifying the Genetics of Disease Inheritance for Bayesian Application

Lawless, D.

2025-03-25 genetic and genomic medicine
10.1101/2025.03.25.25324607 medRxiv
Show abstract

BackgroundAccurate interpretation of genetic variants requires a quantitative estimate of how likely a variant is to contribute to disease, accounting for both observed and unobserved causal alleles across different inheritance modes. MethodsWe developed a statistical framework that computes genome-wide prior probabilities for variant classification by integrating population allele frequencies, disease classifications, and Hardy-Weinberg expectations across dominant, recessive, and X-linked inheritance. Bayesian modelling then combines these priors with individual-level data to produce credible intervals that quantify diagnostic confidence. ResultsThe framework replaces categorical variant classification with continuous posterior probabilities that capture residual uncertainty from incomplete or missing genotype data. Demonstrations in three diagnostic scenarios show accurate quantification of variant-level disease relevance. Application to 557 genes implicated in inborn errors of immunity (IEI) generated a public database of prior probabilities. Integration with protein-protein interaction and immunophenotypic data revealed gene-level constraint patterns, and validation in national cohorts showed close agreement between predicted and observed case numbers. ConclusionsOur method addresses a long-standing gap in clinical genomics by quantifying both observed and unobserved genetic evidence in disease diagnosis. It provides a reproducible probabilistic foundation for variant interpretation, clinical decision-making, and large-scale genomic analysis. 1 AvailabilityThis data is integrated in public panels at https://iei-genetics.github.io. The source code are accessible as part of the variant risk estimation project at https://github.com/DylanLawless/var_risk_est and IEI-genetics project at https://iei-genetics.github.io. The data is available from the Zenodo repository: https://doi.org/10.5281/zenodo.15111583 (Var-RiskEst PanelAppRex ID 398 gene variants.tsv). VarRiskEst is available under the MIT licence. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=199 HEIGHT=200 SRC="FIGDIR/small/25324607v6_ufig1.gif" ALT="Figure 1"> View larger version (52K): org.highwire.dtl.DTLVardef@14e70f2org.highwire.dtl.DTLVardef@d92bd2org.highwire.dtl.DTLVardef@1cbf7d0org.highwire.dtl.DTLVardef@1fa9b76_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.