Back

When the Background Matters: Topic-Dependent reference lists in GWAS and Exome Analyses

Timoney, B.; Guasoni, P.; Zade, K.; Bach, S.; Tropea, D.

2026-08-21 bioinformatics
10.64898/2026.08.14.744838 bioRxiv
Show abstract

Gene Ontology (GO) Biological Process overrepresentation analysis is widely used to interpret gene lists from genetic studies, yet results depend critically on the background (universe/reference list) against which enrichment is tested. This paper examines how genome-exome background mismatch alters GO Biological Process significance and induces annotation-driven bias. First, Monte Carlo simulations across multiple input gene list sizes show that enrichment p-values shift systematically when lists sampled from an exome-like universe are tested against a genome background (and vice versa), producing both inflation and deflation of significance depending on GO term composition; these shifts increase with gene list size. Second, applied analyses of gene lists derived from Genome-Wide Association Studies (GWAS) and Whole Exome Studies (WES) across brain, immune, and metabolic domains demonstrate that background choice changes the set of significant GO IDs, yielding reference-specific terms consistent with both Type I errors (false positives) and Type II errors (false negatives). Because genome backgrounds are commonly used by default, the practical risk is greatest when WES-derived lists are analyzed with genome reference lists. To support reproducible best practice, we provide a simple command set for selecting and documenting study-appropriate backgrounds and for assessing sensitivity of GO Biological Process results to the chosen universe.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.