LRP2: A proteogenomics pipeline for long-read informed protein isoform analysis and discovery
Schertzer, M. D.; Lewandowski, J. T.; Watts, E. F.; Rosenow, W.; Mehlferber, M. M.; Jeffery, E. D.; Adamson, S. I.; Bruand, J.; Tseng, E.; Neelamraju, Y.; Garrett-Bakelman, F. E.; Dolzhenko, E.; Knowles, D. A.; Sheynkman, G.
Show abstract
Most human genes produce multiple RNA isoforms, yet it remains unclear which isoforms are translated into stable, functional proteins. Long-read RNA-sequencing resolves full-length transcript structures and, when paired with mass spectrometry, can provide empirical evidence of isoform translation. Despite this opportunity, comprehensive workflows integrating isoform discovery, open reading frame prediction, peptide identification, and protein inference remain limited, leaving users to handle these steps piecemeal. Here, we present LRP2, a modular, end-to-end long-read proteogenomics pipeline built in Nextflow. LRP2 scales transcript discovery to hundreds of samples via PacBios latest Isocall tool, removes technical artifacts with SQANTI QC, generates and classifies predicted proteomes via CPAT and SQANTI Protein, performs multi-group differential expression and usage analysis via edgeR, DRIMSeq and a long-read adaptation of LeafCutter, and integrates protein-level evidence from DDA and DIA MS data through FragPipe. For cross-dataset comparison of novel isoforms, LRP2 employs deterministic splice-junction, coordinate-based isoform identifiers. Availability and implementationLRP2 is freely available as a modular Nextflow pipeline at: https://github.com/sheynkman-lab/LRP2. LRP2 supports Docker, Apptainer, and Conda environments with GENCODE references. ContactMegan Schertzer, cwp5au@virginia.edu Gloria Sheynkman, gs9yr@virginia.edu
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Eclipse: A Python package for alignment of two or more nontargeted LC-MS metabolomics datasets 95%
- PROTRIDER: Protein abundance outlier detection from mass spectrometry-based proteomics data with a conditional autoencoder 95%
- Generation of ENSEMBL-based proteogenomics databases boosts the identification of non-canonical peptides. 94%
Similar papers in this journal
- Improved open modification searching via unified spectral search with predicted libraries and enhanced vector representations in ANN-SoLo 95%
- Discovery of protein modifications using high resolution differential mass spectrometry proteomics 95%
- prolfqua: A Comprehensive R-package for Proteomics Differential Expression Analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.