Back

A mechanistic model for the negative binomial distribution of single-cell mRNA counts

Amrhein, L.; Harsha, K.; Fuchs, C.

2019-06-03 bioinformatics
10.1101/657619 bioRxiv
Show abstract

Several tools analyze the outcome of single-cell RNA-seq experiments, and they often assume a probability distribution for the observed sequencing counts. It is an open question of which is the most appropriate discrete distribution, not only in terms of model estimation, but also regarding interpretability, complexity and biological plausibility of inherent assumptions. To address the question of interpretability, we investigate mechanistic transcription and degradation models underlying commonly used discrete probability distributions. Known bottom-up approaches infer steady-state probability distributions such as Poisson or Poisson-beta distributions from different underlying transcription-degradation models. By turning this procedure upside down, we show how to infer a corresponding biological model from a given probability distribution, here the negative binomial distribution. Realistic mechanistic models underlying this distributional assumption are unknown so far. Our results indicate that the negative binomial distribution arises as steady-state distribution from a mechanistic model that produces mRNA molecules in bursts. We empirically show that it provides a convenient trade-off between computational complexity and biological simplicity.\n\nGraphical Abstract\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC=\"FIGDIR/small/657619v2_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (38K):\norg.highwire.dtl.DTLVardef@1ba856eorg.highwire.dtl.DTLVardef@8e1b7forg.highwire.dtl.DTLVardef@1af3442org.highwire.dtl.DTLVardef@1901671_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.