Back

When batch correction corrupts gene expression: uncovering distortions in correlation structures

Nourisa, J.; Passemiers, A.; Moreau, Y.; Raimondi, D.

2026-06-10 bioinformatics
10.64898/2026.06.06.729466 bioRxiv
Show abstract

Batch correction is essential for integrating datasets and enabling population-level insights into health and disease. Embedding-based approaches are among the most widely used solutions, but here we highlight a critical, overlooked limitation: these methods can distort feature-to-feature (e.g., gene-gene) relationships, potentially undermining downstream analyses. We investigate this issue and introduce a novel metric to quantify it.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.