A confound-diagnostic toolkit for in silico perturbation with single-cell foundation models
Qiu, R.; Zhao, M. M.
Show abstract
Deleting a gene token from a cells input sequence offers a convenient native strategy for in silico perturbation, but the resulting embedding delta may not represent a biological knockout response. Apparent effects can instead reflect gene identity, universal responsiveness, limited tokenization coverage, library-size contamination, or circular state scoring. Here, we present a confound-diagnostic framework combining held-out increment testing, responsiveness adjustment, coverage gating, library-size diagnostics, and de-circularized state-shift analysis, together with a numerically matched reimplementation of frozen Geneformers perturbation engine. Across Frangieh and Replogle datasets and linear and nonlinear readouts, the native embedding delta provided no reproducible held-out improvement beyond gene identity. Signal-injection calibration showed that the test detected injected residual signal, whereas native increments remained below its detection floor. Matched controls traced apparent positives to raw-count library-size structure, broad responsiveness, and self-referential scoring, while coverage constrained perturbation applicability and estimate stability without establishing biological specificity. This model-adaptable framework helps determine when foundation-model perturbation readouts warrant biological interpretation. MotivationFoundation-model in silico perturbation could predict perturbation effects when matched experimental data are unavailable. However, in zero-shot settings, embedding-derived responses may reflect gene identity, universal responsiveness, tokenization limits, library-size artifacts, or circular state scoring rather than biological knockout effects. We therefore developed a reusable confound-diagnostic framework that applies matched controls to test whether native perturbation readouts contain information beyond these confounds and warrant biological interpretation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Towards inferring causal gene regulatory networks from single cell expression measurements 93%
- Accelerated design of Escherichia coli genomes with reduced size using a whole-cell model and machine learning surrogate 93%
- SCOPE: a normalization and copy number estimation method for single-cell DNA sequencing 92%
Similar papers in this journal
- Robust differential expression testing for single-cell CRISPR screens at low multiplicity of infection 94%
- DiMSum: an error model and pipeline for analyzing deep mutational scanning data and diagnosing common experimental pathologies 93%
- SampleQC: robust multivariate, multi-celltype, multi-sample quality control for single cell data 93%
Similar papers in this journal
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 94%
- StanDep: capturing transcriptomic variability improves context-specific metabolic models 93%
- Capturing cell heterogeneity in representations of cell populations for image-based profiling using contrastive learning 93%
Similar papers in this journal
- Temporal control of sgRNA library activation unlocks large-scale in vivo CRISPR screens 93%
- CRISPRcleanR WebApp: an interactive web application for processing genome-wide pooled CRISPR-Cas9 viability screen 92%
- UniFORM: Towards Universal Immunofluorescence Normalization for Multiplex Tissue Imaging 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.