Back

Prompting Beyond Pairs: Decoupled Semantic Supervision for Knowledge-Guided Multiplex Virtual Staining

Hu, Y.; Wang, J.; Zheng, K.; Yu, H.

2026-07-31 bioengineering
10.64898/2026.07.31.741995 bioRxiv
Show abstract

Virtual staining provides a non-invasive alternative to fluorescence microscopy, yet existing deep learning approaches fundamentally rely on pixel-aligned, multiplexed fluorescence targets for supervision. This dependence on rigidly paired data limits scalability, constrains flexibility in generating diverse subcellular structures, and becomes impractical in data-scarce biological settings. In this work, we introduce a semantic supervision paradigm for virtual staining, demonstrating that domain-knowledge prompts can effectively replace conventional pixel-level supervision. Unlike existing methods constrained by rigidly paired multiplex targets, our framework leverages biological prompts to decouple structural guidance from image translation. This decoupling enables high-fidelity, independent synthesis of multiple subcellular structures using only single-channel data. To ensure high-fidelity generation under weak supervision, we integrate self-supervised representation learning to mitigate data scarcity and incorporate direct preference optimization to suppress structural artifacts. Evaluations on the JUMP benchmark demonstrate that our approach effectively balances flexibility and fidelity, outperforming supervised baselines with a 43.3 % reduction in Average FID and an Average PCC of 0.912, while exhibiting high robustness in channel-deficient scenarios. Furthermore, the model generalizes across four in-house datasets to successfully multiplex six subcellular structures, overcoming the physical constraints of conventional fluorescent staining.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.