Quantifying stability via count splitting to guide model selection in RNA velocity analyses
Li, Y.; Wei, Z. J.; Chen, Y.-C.; Lin, K. Z.
Show abstract
MotivationRNA velocity is a computational framework that predicts future cell states from single-cell RNA sequencing data, offering valuable insights into dynamic biological processes. However, there is a lack of general methods to quantify the uncertainty and stability of these predictions from various RNA velocity methods. ResultsWe present a novel framework for evaluating RNA velocity stability using negative binomial count splitting to generate independent data replicates, a metric we call replicate coherence. Testing five RNA velocity methods across datasets for mouse erythroid, pancreatic, and human brain development, we identify significant performance differences and inconsistencies, such as reversed velocity flows. Our framework remains robust even when intermediary cell states are missing. Furthermore, we introduce a signal-to-random coherence metric to guide model selection. We demonstrate that selecting fits with high replicate coherence uncovers more biologically informative gene pathways. This broadly applicable approach provides a rigorous tool for assessing and comparing RNA velocity methods across diverse biological contexts. Availability and implementationThe code and analyses are available at https://github.com/linnykos/veloUncertainty.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Visualizing scRNA-Seq Data at Population Scale with GloScope 96%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 96%
- Integrating temporal single-cell gene expression modalities for trajectory inference and disease prediction 95%
Similar papers in this journal
- scNODE: Generative Model for Temporal Single Cell Transcriptomic Data Prediction 96%
- A Bayesian framework for inter-cellular information sharing improves dscRNA-seq quantification 95%
- Non-negative Independent Factor Analysis disentangles discrete and continuous sources of variation in scRNA-seq data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.