Back

Spectronaut-nf: A Nextflow Pipeline for Parallel Processing of DIA Data with Spectronaut

Kotimoole, C. N.; Arefian, M.; McKay, E. C.; Kasaragod, S.; Skoraczynski, G.; Collins, B. C.

2026-07-30 bioinformatics
10.64898/2026.07.29.741433 bioRxiv
Show abstract

SummaryContemporary proteomics methods can now generate large-scale DIA datasets of thousands of files that demand substantial computational resources for efficient analysis. Spectronaut is a widely used platform for DIA data processing; however, large-scale searches are often constrained by computational performance and long execution times when run on single workstations. Here, we present Spectronaut-nf, a Nextflow-based pipeline that enables scalable and parallelized execution of Spectronaut analyses across high-performance computing (HPC) environments. The workflow divides directDIA analysis into modular stages, including spectral library generation, DIA searching, and merging results, allowing efficient distribution of tasks across multiple compute nodes. Benchmarking using 72 diaPASEF raw files using typical hardware demonstrated that Spectronaut-nf completed searches in 23.77 hours, compared with 39.09 hours on a Windows workstation and 67.04 hours on a single-node Linux HPC setup. Stress testing with 1,037 diaPASEF raw files further demonstrated the scalability and robustness of the workflow for large proteomics datasets. Across platforms, protein and peptide identifications remained consistent, with only minimal variability attributable to platform-specific differences. Overall, Spectronaut-nf provides a flexible, scalable, and efficient framework for high-throughput DIA proteomics analysis in HPC environments. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=129 SRC="FIGDIR/small/741433v1_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@4a7107org.highwire.dtl.DTLVardef@14287d8org.highwire.dtl.DTLVardef@e484e9org.highwire.dtl.DTLVardef@d21c07_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.