Language of Stains: Tokenization Enhances Multiplex Immunofluorescence and Histology Image Synthesis
Sims, Z.; Govindarajan, S.; Mills, G. B.; Eksi, S. E.; Chang, Y. H.
Show abstract
Multiplex tissue imaging (MTI) is a powerful tool in cancer research, allowing spatially resolved, single-cell phenotype analysis. However, MTI platforms face challenges such as high costs, tissue loss, lengthy acquisition times, and complex analysis of large, multichannel images with batch effects. To address these challenges, we propose a novel computational method to model the interactions between dozens of panel markers and Hematoxylin & Eosin (H&E) staining, enabling in-silico generation of marker stains. This approach reduces the reliance on experimentally measured markers, bridging low-cost H&E data with MTIs high-content information. Our approach uses a two-stage frame-work for channel-wise bioimage synthesis: first, vector quantization learns a visual token vocabulary, then a bidirectional transformer infers missing markers through masked language modeling. Comprehensive bench-marking across different MTI platforms and tissue types demonstrates the effectiveness of our method in improving marker prediction while maintaining biological relevance. This advance makes high-dimensional multiplex tissue imaging more accessible and scalable, supporting deeper insights and potential clinical applications in cancer research.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A comprehensive comparison on cell type composition inference for spatial transcriptomics data 95%
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 95%
- Graph Contrastive Learning of Subcellular-resolution Spatial Transcriptomics Improves Cell Type Annotation and Reveals Critical Molecular Pathways 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.