Back

Deep Learning-Driven Fragment Ion Selection for Improved Quantification in MS based Proteomics

Vu, D. T.; Wallmann, G.; Thielert, M.; Ugur, E.; Oeller, M.; Zwiebel, M.; Ammar, C.; Mann, M.

2026-01-08 bioinformatics
10.64898/2026.01.08.698341 bioRxiv
Show abstract

Quantitative proteomics relies on accurate selection of fragment ions for quantification, yet most current algorithms apply simple strategies such as median intensity or single quality filters. Modern data-independent acquisition (DIA) searches generate rich features such as fragment ion correlations, retention time and many others that could be leveraged to assess fragment quality. We introduce QuantSelect, a novel strategy to select optimal fragments by systematically integrating these features via self-supervised deep learning. QuantSelect uses a regularized, weighted-variance loss on intensity traces normalized via our directLFQ algorithm. This allows learning a fragment quality score without ground truth labels, enabling on-the-fly training on label-free DIA datasets. Integrated within our alphaDIA pipeline, QuantSelect significantly improves quantitative accuracy and in some cases substantially corrects protein intensity estimation. Sensitivity in differential expression improved by 68% in a mixed-species benchmarking dataset and by 18% in single-cell data. QuantSelect provides a practical framework for data-driven fragment selection that improves accuracy, precision and downstream inference in DIA proteomics.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.