Back

Improving inference in wastewater-based epidemiology by modelling the statistical features of digital PCR

Lison, A.; Julian, T.; Stadler, T.

2024-10-17 bioinformatics
10.1101/2024.10.14.618307 bioRxiv
Show abstract

Digital polymerase chain reaction (dPCR) is a powerful technique for quantifying gene targets in environmental samples, with various applications such as biodiversity monitoring and wastewater-based epidemiology. However, statistical analyses of environmental dPCR data often assume, explicitly or implicitly, that concentration measurements have a normal or log-normal error structure, which does not reflect the underlying partitioning statistics of dPCR. Using simulations and real-world environmental data, we show that (log-)normality assumptions are violated for dPCR measurements, leading to inaccurate estimates of gene concentrations and underlying biological processes. To enable reliable analyses of environmental dPCR data, we present a dPCR-specific likelihood model that accounts for concentration-dependent measurement noise and non-detects as characteristic of dPCR assays. We demonstrate that this approach overcomes biases in inference from environmental data, such as estimating free-eDNA decay in seawater or pathogen transmission from wastewater monitoring. Our method is implemented in the R packages "dPCRfit" for regression analyses and "EpiSewer" for wastewater surveillance.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.