Aird-ComboComp: A combinable compressor framework with a dynamic-decider for lossy mass spectrometry data compression
Lu, M.; Tong, J.; Wang, R.; An, S.; Wang, J.; Jiang, H.; Yu, C.
Show abstract
Mass spectrum (MS) data volumes increase with an improved ion acquisition ratio and a highly accurate mass spectrometer. However, the most widely used data format, mzML, does not take advantage of compression methods and improved read performances. Several compression algorithms have been proposed in recent years, and they consider a number of factors, including, numerical precision, metadata read strategies and the compression performance. Due to limited compression ratio, the high-throughput MS data format is still quite large. High bandwidth and memory requirements severely limit the applicability of MS data analysis in cloud and mobile computing. ComboComp is a comprehensive improvement to the Aird data format. Instead of using the general-purpose compressor directly, ComboComp uses two integer-purpose compressors and four general-purpose compressors, and obtains the best compression combination with a dynamic decider, achieving the most balanced compression ratio among all the numerous varieties of compressors. ComboComp supports a seamless extension of the new integer and generic compressors, making it an evolving compression framework. The improvement of compression rate and decoding speed greatly reduces the cost of data exchange and real-time decompression, and effectively reduces the hardware requirements of MS data analysis. Analyzing mass spectrum data on IoT devices can be useful in real-time quality control, decentralized analysis, collaborative auditing, and other scenarios. We tested ComboComp on 11 datasets generated by commonly used MS instruments. Compared with Aird-ZDPD, the compression size can be reduced by an average of 12.9%. The decompression speed is increased by an average of 27.1%. The average compression time is almost the same as that of ZDPD. The high compression rate and decoding speed make the Aird format effective for data analysis on small memory devices. This will enable MS data to be processed normally even on IoT devices in the future. We provide SDKs in three languages, Java, C# and Python, which offer optimized interfaces for the various acquisition modes. All the SDKs can be found on Github: https://github.com/CSi-Studio/Aird-SDK.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- pDeep3: Towards More Accurate Spectrum Prediction with Fast Few-Shot Learning 97%
- mzrtsim: Raw Data Simulation for Reproducible Gas/Liquid Chromatography Mass Spectrometry Based Non-targeted Metabolomics Data Analysis 96%
- spectrum_utils: A Python package for mass spectrometry data processing and visualization 96%
Similar papers in this journal
Similar papers in this journal
- Identification of metabolites from tandem mass spectra with a machine learning approach utilizing structural features 96%
- Alpha-XIC: a deep neural network for scoring the coelution of peak groups improves peptide identification by data-independent acquisition mass spectrometry 95%
- 3DMolMS: Prediction of Tandem Mass Spectra from Three Dimensional Molecular Conformations 95%
Similar papers in this journal
- Deep Learning on Multimodal Chemical and Whole Slide Imaging Data for Predicting Prostate Cancer Directly from Tissue Images 95%
- MealTime-MS: A Machine Learning-Guided Real-Time Mass SpectrometryAnalysis for Protein Identification and Efficient DynamicExclusion 95%
- Effect of terminal phosphate groups on collisional dissociation of RNA oligonucleotide anions 94%
Similar papers in this journal
- Fast alignment of mass spectra in large proteomics datasets, capturing dissimilarities arising from multiple complex modifications of peptides 95%
- Moiety Modeling Framework for Deriving Moiety Abundances from Mass Spectrometry Measured Isotopologues 95%
- MassComp, a lossless compressor for mass spectrometry data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.