AbNovoBench: a resource and benchmarking platform for monoclonal antibody de novo sequencing
Jiang, W.; Luo, L.; Xiong, Y.; Xiao, J.; Lin, Z.; Huang, L.; Zhang, S.; Wang, J.; Wang, C.; Xia, N.; Yuan, Q.; Yu, R.
Show abstract
Monoclonal antibodies (mAbs) are critical in disease diagnostics and therapeutics, yet the performance of mass spectrometry (MS)-based de novo sequencing remains incompletely characterized due to limited antibody-specific datasets and the absence of a standardized benchmark framework. Here we present AbNovoBench, a comprehensive framework for evaluating data analysis strategies for mAb de novo sequencing. It features the largest high-quality dataset to date, generated in-house, comprising 1,638,248 peptide-spectrum matches from 131 mAbs across six species and 11 proteases, supplemented by eight mAbs with known full-length sequence for end-to-end reconstruction assessment. Employing a unified training dataset, we systematically benchmarked 13 deep learning-based de novo peptide sequencing algorithms and three assembly strategies across peptide sequencing metrics (accuracy, robustness, efficiency, error types) and assembly metrics (coverage depth, assembly score). AbNovoBench (https://abnovobench.com) provides an online platform enriched with curated antibody MS resources and pre-trained models, enabling customizable antibody sequencing workflows, accelerating antibody-specific algorithms development, and improving reproducibility in proteomics.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 97%
- mokapot: Fast and flexible semi-supervised learning for peptide detection 96%
- Inserting Pre-Analytical Chromatographic Priming Runs Significantly Improves Targeted Pathway Proteomics With Sample Multiplexing 96%
Similar papers in this journal
Similar papers in this journal
- Carafe enables high quality in silico spectral library generation for data-independent acquisition proteomics 97%
- FIDDLE: a deep learning method for chemical formulas prediction from tandem mass spectra 96%
- SugarQuant: a streamlined pipeline for multiplexed quantitative site-specific N-glycoproteomics 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.