Back

Identifying maximally informative signal-aware representations of single-cell data using the Information Bottleneck

Dubnov, S.; Piran, Z.; Soreq, H.; Nitzan, M.

2024-07-17 bioinformatics
10.1101/2024.05.22.595292 bioRxiv
Show abstract

Rapid advancements in single-cell RNA-sequencing (scRNA-seq) technologies revealed the richness of myriad attributes encompassing cell identity. However, the complexity of the data hinders tasks focusing on a specific biological signal. To address this challenge, we introduce bioIB, a framework based on the Information Bottleneck method, designed to extract an interpretable compressed representation of scRNA-seq data, optimally-informative with respect to a desired biological signal, such as developmental stage or disease state. Provided with cellular labels representing the signal of interest, bioIB generates weighted gene clusters, termed metagenes, that compress the data, while maximizing signal-specific information. Following the Information Bottleneck principle, bioIB identifies an optimal trade-off between data compression and retaining target information. Further, bioIB provides the hierarchical structure of the metagenes, revealing the interconnections between the corresponding biological processes and cellular populations, such as the developmental hierarchy of hematopoietic cell types. We showcase bioIBs applicability to diverse biological contexts, including Alzheimers Disease, epithelial-to-mesenchymal transition, immune development and hematopoiesis, demonstrating that the compressed representations capture signal-associated molecular pathways and expose cellular subpopulations with prominent phenotypes such as transition states and disease association.

Published in Cell Systems (predicted rank #8) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.