Back

TREVI: A Transcriptional Regulation-driven Variational Inference Model to Speculate Gene Expression Mechanism with Integration of Single-cell Multi-omics

Cao, L.; Zhang, W.; Zeng, F.; Wang, Y.

2023-12-18 bioinformatics
10.1101/2023.11.22.568363 bioRxiv
Show abstract

Single-cell multi-omics technology enables the concurrent measurement of multiple molecular entities, making it critical for unraveling the inherent gene regulation mechanisms driving cell heterogeneity. However, existing multi-omics techniques have limitations in capturing the intricate regulatory interactions among these molecular components. In this study, we introduce TREVIXMBD (Transcriptional REgulation-driven Variational Inference), a novel method that integrates the well-established gene regulation structure with scRNA-seq and scATAC-seq data through an advanced Bayesian framework. TREVIXMBD models the generation of gene expression profiles in individual cells by considering the integrated influence of three fundamental biological factors: accessibility of cis-regulatory elements regions, transcription factor (TF) activities and regulatory weights. TF activities and regulatory weights are probabilistically represented as latent variables, which capture the inherent gene regulatory significance. Hence, in contrast to gene expression, TF activities and regulatory weights that depict the cell states from a more intrinsic perspective, can keep consistent across diverse datasets. TREVIXMBD exhibits superior performance when compared to baseline methods in a variety of biological analyses, including cell typing, cell development tracking, and batch effect correction, as validated through comprehensive benchmarking. Moreover, TREVIXMBD can reveal variations in TF-gene regulation relationships across cells. The pretrained TREVIXMBD model can work even when only scRNA-seq is available. Overall, TREVIXMBD introduces a pioneering biological-mechanism-driven framework for elucidating cell states at a gene regulatory level. The models structure is adaptable for the inclusion of additional biological factors, allowing for flexible and more comprehensive gene regulation analysis.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.