Back

A Pan-Organ Vision-Language Model for Generalizable 3D CT Representations

Beeche, C. A.; Kim, J.; Tavolinejad, H.; Zhao, B.; Sharma, R.; Duda, J.; Gee, J.; Dako, F.; Verma, A.; Morse, C.; Hou, B.; Shen, L.; Sagreiya, H.; Davatzikos, C.; Damrauer, S. M.; Ritchie, M. D.; Rader, D. J.; Long, Q.; Chen, T.; Kahn, C. E.; Chirinos, J. A.; Witschey, W. R.; Penn Medicine Biobank,

2025-07-03 radiology and imaging
10.1101/2025.07.03.25330654 medRxiv
Show abstract

Vision-language foundation models (VLMs) for computed tomography (CT) are emerging tools capable of learning generalizable representations from large-scale clinical imaging data. Yet, it remains unclear to what extent these models encode biologically meaningful information relevant to real-world clinical variation. We introduce Percival, a CT-native VLM trained on more than 400,000 CT-report pairs from the Penn Medicine BioBank using a dual-encoder symmetric contrastive framework, with the objective of characterizing the biological associations embedded through contrastive pretraining. Across over 20,000 held-out participants, Percivals latent space shows strong alignment with clinical attributes, body-size measures, and multiple laboratory biomarkers. Phenome-wide analyses further reveal broad correspondence between latent features and disease phenotypes, including conditions not typically evaluated by CT; survival analyses demonstrate that the embeddings capture longitudinal risk patterns. Together, these findings reveal that CT-VLMs uncover a rich latent structure aligned with physiological measurements and disease phenotypes spanning the disease-prevalence spectrum.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.