MatchCLOT: Single-Cell Modality Matching with Contrastive Learning and Optimal Transport
Gossi, F.; Pati, P.; Martinelli, A. L.; Rapsomaniki, M. A.
Show abstract
Recent advances in single-cell technologies have enabled the simultaneous quantification of multiple biomolecules in the same cell, opening new avenues for understanding cellular complexity and heterogeneity. However, the resulting multimodal single-cell datasets present unique challenges arising from the high dimensionality of the data and the multiple sources of acquisition noise. In this work, we propose MO_SCPLOWATCHC_SCPLOWCLOT, a novel method for single-cell data integration based on ideas borrowed from contrastive learning, optimal transport, and transductive learning. In particular, we use contrastive learning to learn a common representation between two modalities and apply entropic optimal transport as an approximate maximum weight bipartite matching algorithm. Our model obtains state-of-the-art performance in the modality matching task from the NeurIPS 2021 multimodal single-cell data integration challenge, improving the previous best competition score by 28.9%. Our code can be accessed at https://github.com/AI4SCR/MatchCLOT.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Graph Contrastive Learning of Subcellular-resolution Spatial Transcriptomics Improves Cell Type Annotation and Reveals Critical Molecular Pathways 96%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 96%
- Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing data 96%
Similar papers in this journal
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 96%
- scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis 95%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.