Back

PeakPerformance - A software tool for fitting LC-MS/MS peaks including uncertainty quantification by Bayesian inference

Niesser, J.; Osthege, M.; von Lieres, E.; Wiechert, W.; Noack, S.

2024-02-21 systems biology
10.1101/2024.02.17.580815 bioRxiv
Show abstract

A major bottleneck of chromatography-based analytics has been the elusive fully automated identification and integration of peak data without the need of extensive human supervision. The presented Python package PeakPerformance applies Bayesian inference to chromatographic peak fitting, and provides an automated approach featuring model selection and uncertainty quantification. Currently, its application is focused on data from targeted liquid chromatography tandem mass spectrometry (LC-MS/MS), but its design allows for an expansion to other chromatographic techniques. PeakPerformance is implemented in Python and the source code is available on GitHub. It is unit-tested on Linux and Windows and accompanied by general introductory documentation, as well as example notebooks. Practical applicationThe presented PeakPerformance tool performs automated chromatographic peak data fitting using Bayesian methodology. Accordingly, it innovates by delivering built-in uncertainty quantification for each peak, thus taking the measurement noise into account. Using a convergence statistic and based on the determined peak uncertainties, the differentiation of signals into peak and noise was improved and false positives or negatives were largely eliminated. The provided documentation and the implemented convenience functions are meant to lower the barrier of entry for users with little programming experience. Lastly, the modular design of the software enables modification and expansion to data from different chromatographic methods.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.