Back

scMoE: single-cell Multi-Modal Multi-Task Learning via Sparse Mixture-of-Experts

Yun, S.; Peng, J.; Lee, N.; Zhang, Y.; Park, C.; Liu, Z.; Chen, T.

2024-11-15 bioinformatics
10.1101/2024.11.12.623336 bioRxiv
Show abstract

Recent advances in measuring high-dimensional modalities, including protein levels and DNA accessibility, at the single-cell level have prompted the need for frameworks capable of handling multi-modal data while simultaneously addressing multiple tasks. Despite these advancements, much of the work in the single-cell domain remains limited, often focusing on either a single-modal or single-task perspective. A few recent studies have ventured into multimodal, multi-task learning, but we identified a [circled1] Optimization Conflict issue, leading to suboptimal results when integrating additional modalities, which is undesirable. Furthermore, there is a [circled2] Costly Interpretability challenge, as current approaches predominantly rely on costly post-hoc methods like SHAP. Motivated by these challenges, we introduce scMoE1, a novel framework that, for the first time, applies Sparse Mixture-of-Experts (SMoE) within the single-cell domain. This is achieved by incorporating an SMoE layer into a transformer block with a cross-attention module. Thanks to its design, scMoE inherently possesses mechanistic interpretability, a critical aspect for understanding underlying mechanisms when handling biological data. Furthermore, from a post-hoc perspective, we enhance interpretability by extending the concept of activation vectors (CAVs). Extensive experiments on simulated datasets, such as Dyngen, and real-world multi-modal single-cell datasets, including {DBiT-seq, Patch-seq, ATAC-seq}, demonstrate the effectiveness of scMoE. Source code of scMoE is available at: https://github.com/UNITES-Lab/scMoE.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.