Benchmarking feature quality assurance strategies for non-targeted metabolomics
El Abiead, Y.; Milford, M.; Schoeny, H.; Rusz, M.; Salek, R. M.; Koellensperger, G.
Show abstract
Automated data pre-processing (DPP) forms the basis of any liquid chromatography-high resolution mass spectrometry-driven non-targeted metabolomics experiment. However, current strategies for quality control of this important step have rarely been investigated or even discussed. We exemplified how reliable benchmark peak lists could be generated for eleven publicly available datasets acquired across different instrumental platforms. Moreover, we demonstrated how these benchmarks can be utilized to derive performance metrics for DPP and tested whether these metrics can be generalized for entire datasets. Relying on this principle, we cross-validated different strategies for quality assurance of DPP, including manual parameter adjustment, variance of replicate injection-based metrics, unsupervised clustering performance, automated parameter optimization, and deep learning-based classification of chromatographic peaks. Overall, we want to highlight the importance of assessing DPP performance on a regular basis.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Boosting the MS1-only proteomics with machine learning allows 2000 protein identifications in 5-minute proteome analysis 96%
- Data-Driven Optimization of DIA Mass Spectrometry by DO-MS 96%
- Hybrid Quadrupole Mass Filter Radial Ejection Linear Ion Trap and Intelligent Data Acquisition Enable Highly Multiplex Targeted Proteomics 96%
Similar papers in this journal
- WiPP: Workflow for improved Peak Picking for Gas Chromatography-Mass Spectrometry (GC-MS) data 96%
- Scan-Centric, Frequency-Based Method for Characterizing Peaks from Direct Injection Fourier transform Mass Spectrometry Experiments 96%
- Robust Moiety Model Selection Using Mass Spectrometry Measured Isotopologues 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.