Back

Phenotype Prediction using a Tensor Representation and Deep Learning from Data Independent Acquisition Mass Spectrometry

Zhang, F.; Yu, S.; Wu, L.; Zang, Z.; Yi, X.; Zhu, J.; Lu, C.; Sun, P.; Sun, Y.; Selvarajan, S.; Chen, L.; Teng, X.; Zhao, Y.; Wang, G.; Xiao, J.; Huang, S.; Kon, O. L.; Iyer, G. N.; Li, S. Z.; Luan, Z.; Guo, T.

2020-03-06 bioinformatics
10.1101/2020.03.05.978635 bioRxiv
Show abstract

A novel approach for phenotype prediction is developed for mass spectrometric data. First, the data-independent acquisition (DIA) mass spectrometric data is converted into a novel file format called "DIA tensor" (DIAT) which contains all the peptide precursors and fragments information and can be used for convenient DIA visualization. The DIAT format is fed directly into a deep neural network to predict phenotypes without the need to identify peptides or proteins. We applied this strategy to a collection of 102 hepatocellular carcinoma samples and achieved an accuracy of 96.8% in classifying malignant from benign samples. We further applied refined model to 492 samples of thyroid nodules to predict thyroid cancer; and achieved a predictive accuracy of 91.7% in an independent cohort of 216 test samples. In conclusion, DIA tensor enables facile 2D visualization of DIA proteomics data as well as being a new approach for phenotype prediction directly from DIA-MS data.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.