Back

Predicting single-cell responses to novel genetic perturbations with optimal transport

Do, C.; Lähdesmäki, H.

2026-01-08 bioinformatics
10.64898/2026.01.08.698172 bioRxiv
Show abstract

The combination of pooled genetic perturbation screening with single-cell RNA-sequencing has allowed for large-scale profiling of single-cell transcriptional responses, providing valuable insights into gene function and cellular behavior. There is a need for computational methods capable of predicting single-cell responses to novel genetic perturbations, thereby allowing for the efficient exploration of the vast combinatorial space of potential multi-gene perturbations. Optimal transport (OT) presents a compelling framework to model this type of data given the lack of true control-perturbed cell pairings. However, existing neural OT models are generally constrained to the perturbations already seen in training, while the use of OT theory in methods capable of modeling novel genetic perturbations remains limited. To bridge this gap, we present an OT-based method to predict single-cell transcriptional responses to novel genetic perturbations. Leveraging a flexible attention-based aggregation approach, our method integrates gene representations from multiple sources of prior knowledge, spanning from functional descriptions of genes to gene-gene relationships, producing a biologically informed representation of the perturbation. Moreover, we utilize an OT-based loss to align predicted and observed perturbed cell populations, avoiding the need to assign random control-perturbed cell pairings. Experiments on three benchmark datasets demonstrate the highly competitive performance of our model compared to state-of-the-art methods. Further evaluations on differential expression analysis and genetic interaction modeling demonstrate the biological relevance and potential utility of our model across diverse applications.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.