Back

A data-driven consensus framework for Ct interpretation in real-world multi-assay qPCR diagnostics

Wang, J.; Chen, J.; Zhao, B.; Zhang, G.; Jian, S.; Deng, T.; Liang, D.

2026-06-15 pathology
10.64898/2026.06.11.26355491 medRxiv
Show abstract

While cycle threshold (Ct) values from quantitative PCR (qPCR) serve as the gold-standard indicators of target abundance, their clinical interpretation is frequently confounded by inherent variability across diverse assay designs, reagents, and instrumentation. In this study, we present a data-driven consensus framework for Ct evaluation that uses large-scale, multi-assay amplification data to establish reference patterns of normal Ct behavior. Based on a total of 41,770 amplification curves collected from four routine diagnostic assays across two PCR platforms, we evaluated machine learning models across three experimental scenarios: within-platform validation, cross-assay generalization, and cross-platform transfer. Extreme gradient boosting (XGBoost) achieved the most accurate and stable predictions under data-sufficient, within-platform conditions with a mean absolute error (MAE) of 0.0419, while pooled multi-assay training improved cross-assay robustness compared with single-assay models. Model performance was further assessed using a deviation-based metric to quantify differences between predicted and instrument-reported Ct values, allowing efficient identification of anomalous amplification curves in large datasets. Notably, direct application across platforms without recalibration led to a substantial decline in performance, with a MAE of 2.62, showing platform-dependent variability. These findings indicate strong stability under within-platform and cross-assay conditions, with scalability contingent upon appropriate cross-platform calibration.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.