Back

A procedure for controlling the false discovery rate of de novopeptide sequencing

Sanders, J.; Noble, W. S.; Keich, U. S.

2025-09-17 bioinformatics
10.1101/2025.09.12.675837 bioRxiv
Show abstract

De novo sequencing is a powerful method for identifying peptides from mass spectrometry proteomics experiments without the use of a protein database. However, applications of de novo sequencing are currently severely limited by the lack of a reliable procedure for controlling the false discovery rate (FDR). Here, we introduce an FDR control procedure for the de novo setting which is at least as powerful as database search, give empirical evidence that it is statistically valid, and demonstrate its utility on a set of common de novo applications.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.