Back

PCRedux: A Data Mining and Machine Learning Toolkit for qPCR Experiments

Burdukiewicz, M.; Spiess, A.-N.; Rafacz, D.; Blagodatskikh, K.; Huggett, J.; McCall, M. N.; Schierack, P.; Rödiger, S.

2021-04-01 bioinformatics
10.1101/2021.03.31.437921 bioRxiv
Show abstract

MotivationQuantitative Real-time PCR (qPCR) is a widely used -omics method for the precise quantification of nucleic acids, in which the result is associated with the presence/absence or quantity of a specific nucleic acid sequence. As the amount of qPCR data increases worldwide, the manual assessment of results becomes challenging and difficult to reproduce. To overcome this, some automatable characteristics of amplification curves have been described in the literature, often with an appropriate "rule of thumb". ResultsWe developed PCRedux to analyze and calculate 90 numerical qPCR amplification curve descriptors ( features") from large datasets of qPCR amplification curves that are aimed for interpretable machine learning and development of decision support systems. In a case study of a diverse dataset with 3181 positive, negative and ambiguous amplification curves, as assessed by three human raters, we demonstrate a sensitivity >99 % and specificity >97 % in detecting positive and negative amplification. PCRedux is unique as it goes beyond traditional qPCR analysis to capture curvature properties that improve the characterization and classification of amplification curves. The calculation of the features is reproducible and objective, since R is used as a controllable working environment. PCRedux is not a black box, but open source software following on the principle of mathematically interpretable features. These can be combined with user-defined labels for automatic multi-category classification and regression in machine learning. Availabilityhttps://cran.r-project.org/package=PCRedux. Web server: http://shtest.evrogen.net/PCRedux-app/. Documentation: https://PCRuniversum.github.io/PCRedux/.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.