AlphaVaR: an R framework for the statistical interpretation of AlphaGenome variant-effect predictions
Marhaba, K.; Maj, C.; Schumacher, J.; Dasmeh, P.
Show abstract
SummaryAlphaGenome (Google DeepMind) scores a DNA variant across thousands of functional tracks at single-base resolution, reporting both the magnitude of each predicted effect and its rarity against a genome-wide background. That volume is itself the obstacle to biological interpretation. Here we present AlphaVaR, an R package that gives AlphaGenomes output a typed structure together with the statistical methods and visualizations needed to interpret it. The output schema is identical for every variant, so the same tests apply throughout it. AlphaVaR provides localization tests with multiple-testing correction and effect sizes, a specificity index measuring how far an effect concentrates on a few elements of a chosen variable, and a transparent prioritization that ranks candidates across interpretable criteria and maps each to a target gene. Results feed a plot library, reproducible reports and a code-free Shiny application. Applied to rs1427407, the lead common variant for fetal-haemoglobin level, AlphaVaR recovers the established biology of the BCL11A erythroid enhancer. Availability and implementationhttps://github.com/KarimMarhaba/AlphaVaR, released under the MIT licence, R [≥] 4.2, with documentation at https://karimmarhaba.github.io/AlphaVaR/. The released version is archived at Zenodo (doi:10.5281/zenodo.21939265); the AlphaGenome scores analysed here are archived as a separate dataset (doi:10.5281/zenodo.21920988), and the scripts that regenerate every figure and reported number are in the repository (Supplementary Section S5). Contactpouria.dasmeh@uni-marburg.de Supplementary informationSupplementary data are available online.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.