Reconstructing developmental and disease progression with sample-level embeddings
Jiang, L.; Tan, Z. C.; Grabski, I. N.; Hao, Y.; Nakatsuka, N.; Sarkar, S.; Shenoy, A.; Satija, R.
Show abstract
Single-cell genomics is transformative for characterizing cellular heterogeneity, but many translational questions require comparing samples rather than individual cells. Typical analyses compare "case" and "control" groups, ignoring sample-level variation within them. Here we present scSLIDE, a framework that transforms each samples single-cell data into a compact profile describing where its cells fall in high-dimensional space. By comparing these density profiles, scSLIDE calculates embeddings to characterization variation across samples. Applied to COVID-19 infection, Alzheimers disease, and zebrafish embryogenesis, we show that scSLIDE can be used to cluster patients, reconstruct sample-level disease trajectories, and identify coordinated cellular programs across samples. We discover independent axes of infection and severity, a molecular disease progression that aligns with pathology-based estimates of neurodegeneration, and embryonic "pseudostages" varying across and within timepoints. In each case, we demonstrate how case-control analyses collapse rich biological variation into binary labels, and how sample-level embedding represents a powerful analytical alternative.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.