The Everything Bagel Feature Finder: Ultra-fast automated feature finding for untargeted metabolomics
Shin, Y.; El Abiead, Y.; Jarmusch, A. K.; Strobel, M.; Abraham, P. E.; Thurmon, S.; Acharya, D. D.; Aron, A.; Bilbao, A.; Bowen, B. P.; Broeckling, C. D.; Brown, C. J.; Charron-Lamoureux, V.; Chen, X.; Damiani, T.; Doty, A.; Du, X.; Garg, N.; Papadopoulos Lambidis, S.; McCall, L.-I.; Kirkwood-Donelson, K. I.; Northen, T.; Prenni, J.; Rennie, E. E.; Vining, O. B.; Wang, C. X.; Xiong, Q.; Zhao, H. N.; Dorrestein, P. C.; Petras, D.; Phelan, V. V.; Wang, M.
Show abstract
Metabolomics studies are increasingly being applied with hundreds to thousands, even tens of thousands of samples that demand rapid, automated data processing while maintaining analytical sensitivity or quantitative accuracy. A major computational bottleneck is feature finding, which is the transformation of LC-MS and LC-MS/MS data into a set of analyte signals aligned and quantified across samples. Feature finding can be computationally intensive and often requires manual iterative parameter optimization. To accelerate this process, we present the Everything Bagel (EB) feature finder, an ultra-fast automated feature finding tool that integrates feature detection, retention-time alignment, and gap filling designed for run-time and memory efficiency. We benchmarked EB against two automated feature finding methods on eight benchmarking datasets. Specifically, we evaluated these three feature finding methods by measuring spike-in standard detection coverage, dilution series quantification accuracy, and yeast 12C/13C credentialed features. In this evaluation, the EB feature finder achieved performance comparable to, and often exceeding, existing methods while requiring up to 150-fold lower CPU hours and up to 113-fold lower wall time. We further demonstrated the bioanalytical validity of EB by reanalyzing published datasets used for biomarker discovery and reproduced biologically significant features that matched the published findings using manually tuned feature finding settings. Taken along with the speed improvements, we anticipate EB will enhance the ability to automatically analyze datasets with thousands to tens of thousands of samples for the community.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Concepts and Software Package for Efficient Quality Control in Targeted Metabolomics Studies - MeTaQuaC 97%
- SmartPeak automates targeted and quantitative metabolomics data processing 96%
- Met-ID: An Open-Source Software for Comprehensive Annotation of Multiple On-Tissue Chemical Modifications in MALDI-MSI 96%
Similar papers in this journal
- WiPP: Workflow for improved Peak Picking for Gas Chromatography-Mass Spectrometry (GC-MS) data 96%
- Scan-Centric, Frequency-Based Method for Characterizing Peaks from Direct Injection Fourier transform Mass Spectrometry Experiments 95%
- Information-Content-Informed Kendall-tau Correlation: Utilizing Missing Values 95%
Similar papers in this journal
- Evaluation of a prototype Orbitrap Astral Zoom mass spectrometer for quantitative proteomics - Beyond identification lists 96%
- Low-temperature HILIC provides enhanced separations and stability for LC-MS-based metabolomics 96%
- Hybrid Quadrupole Mass Filter Radial Ejection Linear Ion Trap and Intelligent Data Acquisition Enable Highly Multiplex Targeted Proteomics 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.