Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling
Burq, M.; Stepec, D.; Kim, C.; Cimermancic, P.
Show abstract
Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 93%
- Clinical classifiers of COVID-19 infection from novel ultra-high-throughput proteomics 93%
- Cell shapes decode molecular phenotypes in image-basedspatial proteomics 92%
Similar papers in this journal
Similar papers in this journal
- An adaptive, continuous-learning framework for clinical decision-making from proteome-wide biofluid data 97%
- LEOPARD: missing view completion for multi-timepoints omics data via representation disentanglement and temporal knowledge transfer 96%
- CAPTAIN: A multimodal foundation model pretrained on co-assayed single-cell RNA and protein 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.