Back

De Novo sequencing-assisted homology search for DIA data analysis enables low abundance peptide variants discovery

Qiao, R.; Li, H.; Bian, H.; Xin, L.; Shan, B.

2025-06-02 bioinformatics
10.1101/2025.05.30.657054 bioRxiv
Show abstract

Data-independent acquisition mass spectrometry (DIA-MS) has emerged as a powerful approach for comprehensive proteome profiling. Spectral library search and library-free search are the two major approaches for DIA data analysis. The spectral library search requires high-quality spectral libraries derived from the search results of data-dependent acquisition (DDA) experiments, while library-free approaches rely on prediction models to generate in silico libraries. Both methodologies constrain the search space to the peptide list in the database, limiting the discovery of variant peptides arising from genetic variations or mutations. We present a novel computational method DIAVariant designed to identify peptide sequence variants directly and solely from complex DIA spectra while rigorously controlling the false discovery rate. Our experimental results demonstrate that DIAVariant successfully identifies sequence variants previously detected through proteogenomic approaches, while maintaining high specificity across multiple datasets. When integrated with existing DIA database search solutions, our approach constitutes a comprehensive analytical workflow capable of identifying peptides both represented within reference protein databases and those arising from sequence variations not captured in standard databases.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.