Back

Interpretable models for scRNA-seq data embedding with multi-scale structure preservation

Novak, D.; de Bodt, C.; Lambert, P.; Lee, J. A.; Van Gassen, S.; Saeys, Y.

2024-10-03 bioinformatics
10.1101/2023.11.23.568428 bioRxiv
Show abstract

The ability to explore high-dimensional single-cell transcriptomics data efficiently is crucial in many biological studies. Dimensionality reduction techniques have therefore emerged as a basic building block of analytical work-flows. They generate low-dimensional embeddings that capture important structures in the data, and are often used in discovery, quality control, and downstream analysis. However, the trustworthiness of current methods and the rigour of popular evaluation criteria are limited. We tackle this in an empirical study of structure-preserving data embeddings, delivering two new tools. First, we introduce ViScore: a robust scoring framework that improves both unsupervised and supervised quality metrics, with emphasis on scalability and fairness. Second, we introduce ViVAE: a deep learning model that achieves better multi-scale structure preservation and is equipped with new tools for interpretability. We demonstrate the potential of our framework to advance the trustworthiness of single-cell dimensionality reduction in a quantitative comparison and focused case studies. ContextO_LIThis paper advances empirical evaluation of dimensionality reduction (DR) in single-cell transcriptomics and proposes a new DR model that achieves a favourable balance between local and global structure preservation. Both the evaluation methodology and the DR model are readily applicable to a broad range of single-cell data. C_LIO_LIEvaluation of single-cell DR methods has relied on very informal definitions of local and global structures. There has been an emphasis on evaluation metrics that are straightforward to use in benchmarks involving representative datasets (57; 65; 24). This has been in lieu of extensive descriptions of the underlying data manifolds, their intrinsic dimensionality, and dataset-specific topologies, which is infeasible with much of single-cell omics data. C_LIO_LIMathematical descriptions of multi-scale structures in high-dimensional point clouds are possible within the topology domain (1). However, applying them to real biological data is difficult and can require considerable supervision and tuning (21). For this reason, the present study focuses on broadly applicable evaluation strategies that make few assumptions about the datasets at hand. C_LIO_LIThe proposed DR model is shown to separate distinct cell compartments and represent continuities and discontinuities well, based on case studies with both single-snapshot and developmental single-cell transcripmtomics data. C_LIO_LIAlthough some DR methods are equipped with quality-control and explainability measures (14; 35), their general adoption has been limited. Encoder indicatrices, a tool for visualising unwanted embedding distortions, are presented here to advance the explainability of single-cell data embeddings. C_LIO_LIThe new scoring framework (ViScore) and DR model (ViVAE) are published as Python packages, along with reproducible experiments, guided tutorials, and an automated benchmarking framework. Individual elements of both can be integrated into alternative models, contributing to advancement in the single-cell DR field. C_LI

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.