Towards CNN Representations for Small Mass Spectrometry Data Classification: From Transfer Learning to Cumulative Learning
Seddiki, K.; Saudemont, P.; Precioso, F.; Ogrinc, N.; Wisztorski, M.; Salzet, M.; Fournier, I.; Droit, A.
Show abstract
Rapid and accurate clinical diagnosis of pathological conditions remains highly challenging. A very important component of diagnosis tool development is the design of effective classification models with Mass spectrometry (MS) data. Some popular Machine Learning (ML) approaches have been investigated for this purpose but these ML models require time-consuming preprocessing steps such as baseline correction, denoising, and spectrum alignment to remove non-sample-related data artifacts. They also depend on the tedious extraction of handcrafted features, making them unsuitable for rapid analysis. Convolutional Neural Networks (CNNs) have been found to perform well under such circumstances since they can learn efficient representations from raw data without the need for costly preprocessing. However, their effectiveness drastically decreases when the number of available training samples is small, which is a common situation in medical applications. Transfer learning strategies extend an accurate representation model learnt usually on a large dataset containing many categories, to a smaller dataset with far fewer categories. In this study, we first investigate transfer learning on a 1D-CNN we have designed to classify MS data, then we develop a new representation learning method when transfer learning is not powerful enough, as in cases of low-resolution or data heterogeneity. What we propose is to train the same model through several classification tasks over various small datasets in order to accumulate generic knowledge of what MS data are, in the resulting representation. By using rat brain data as the initial training dataset, a representation learning approach can have a classification accuracy exceeding 98% for canine sarcoma cancer cells, human ovarian cancer serums, and pathogenic microorganism biotypes in 1D clinical datasets. We show for the first time the use of cumulative representation learning using datasets generated in different biological contexts, on different organisms, in different mass ranges, with different MS ionization sources, and acquired by different instruments at different resolutions. Our approach thus proposes a promising strategy for improving MS data classification accuracy when only small numbers of samples are available as a prospective cohort. The principles demonstrated in this work could even be beneficial to other domains (astronomy, archaeology...) where training samples are scarce.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Extremely fast and accurate open modification spectral library searching of high-resolution mass spectra using feature hashing and graphics processing units 96%
- Optimizing accuracy and depth of protein quantification in experiments using isobaric carriers 96%
- PAMPA: a software for peptide markers and taxonomic identification for ZooMS samples in Archaeology and Paleontology 96%
Similar papers in this journal
- Deep Learning on Multimodal Chemical and Whole Slide Imaging Data for Predicting Prostate Cancer Directly from Tissue Images 97%
- MealTime-MS: A Machine Learning-Guided Real-Time Mass SpectrometryAnalysis for Protein Identification and Efficient DynamicExclusion 96%
- Ultraviolet Photodissociation of tryptic peptide backbones at 213 nm 96%
Similar papers in this journal
Similar papers in this journal
- Towards robust machine olfaction: debiasing GC-MS data enhances prostate cancer diagnosis from urine volatiles 96%
- XFlow: An algorithm for extracting ion chromatograms 96%
- Interpretable dimensionality reduction and classification of mass spectrometry imaging data in a visceral pain model via non-negative matrix factorization 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.