Back

Unambiguous signatures of malignancies extracted from images of growing cells

Kalweit, G.; Kalweit, M.; Checinska, W.; Saric, M.; Berger, R.; Bodurova-Spassova, E.; Rawluk, J.; Talvard-Balland, N.; Klett, A.; Follo, M.; Kreutmair, S.; Duque-Afonso, J.; Lübbert, M.; Zeiser, R.; Frank, J.; Mertelsmann, R.

2026-01-13 oncology
10.64898/2026.01.10.26343803 medRxiv
Show abstract

As malignant transformation arises from dysregulated cellular programming with characteristic morphological changes, we hypothesized that cancer cells exhibit a unique morphological signature detectable in microscopy images of cells in vitro and in situ across modalities. To test this, we developed CellSign, an AI-based framework to generate Cell Dynamics Fingerprints, which (a) reconstruct morphological progression, (b) remove physiological variation, and (c) support automated malignancy assessment. To arrive there, cell morphology is represented by embeddings from vision foundation models, and a process we term Healthy-Component Reduction is used to refine these by subtracting normal physiological variation, thereby exposing residual disease-specific cues. Embeddings from healthy and malignant cells are organized with manifold learning and summarized with kernel density estimation. We show that unambiguous malignant signatures exist and that our method is robust across diverse datasets spanning breast cancer, lung cancer, and leukemia, transferring reliably between populations of single-cell images and multi-cell patches. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=179 SRC="FIGDIR/small/26343803v1_ufig1.gif" ALT="Figure 1"> View larger version (60K): org.highwire.dtl.DTLVardef@61ee1borg.highwire.dtl.DTLVardef@15735fdorg.highwire.dtl.DTLVardef@999587org.highwire.dtl.DTLVardef@1280ee5_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.