Back

COMPASS: Component-Wise Inference of Shared and Gene-Specific Perturbation Response

Liang, H.; Singh, R.

2026-08-06 bioinformatics
10.64898/2026.08.03.742643 bioRxiv
Show abstract

Predicting how a genetic perturbation reshapes a cells transcriptome is a central goal of computational biology. Previous studies report that the mean response across training perturbations rivals specialized models on standard accuracy metrics, even though it cannot distinguish which perturbation occurred. Across 2,270 CRISPRi perturbations measured in each of six cell lines, we show that this apparent paradox reflects a conserved organization of perturbation responses. Perturbations span a continuum from responses strongly aligned with the mean to more targeted responses that depart from it. Crucially, a perturbations position along this continuum is conserved across cell lines (Kendalls W = 0.59) and predictable from STRING protein-interaction embeddings (R2 = 0.35). We formalize this structure with COMPASS, an interpretable linear model that decomposes each response into shared and gene-specific components and estimates them separately. The shared-response component is modeled as a cell-line-wide response scaled by a perturbation-specific coefficient. This coefficient is strongly conserved across cell lines. The residual gene-specific component--which is moderately conserved across cell lines--recovers pathway-level programs. COMPASS outperforms scGPT, CPA, GEARS, GenePert, and SO_SCPLOWTATEC_SCPLOW in both response accuracy (de-biased Pearson delta 0.34 vs. [≤] 0.32) and perturbation discrimination (cosine PDS gain 0.23 vs. [≤] 0.08). These results recast perturbation prediction across cellular contexts as component-wise inference, with each component estimated from the evidence best suited to it.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.