Sample-specific haplotype-resolved protein isoform characterization via long-read RNA-seq-based proteogenomics
Wissel, D.; Sheynkman, G. M.; Robinson, M. D.
Show abstract
Protein isoform inference from bottom-up mass spectrometry (MS) relies on database search strategies that assume the reference protein database accurately reflects the full repertoire of genetic and transcriptomic states present in the sample being analyzed. Long-read RNA sequencing (lrRNA-seq) now enables simultaneous recovery of complete transcript (splice) structures and the genetic variants present on each molecule, offering a direct route to allele-specific isoforms. Yet, this capability has not been fully leveraged to improve MS-based proteogenomics workflows. Here, we develop an end-to-end workflow for constructing and searching haplotype-resolved, sample-specific proteomes using matched lrRNA-seq and MS data. We benchmark phasing algorithms on PacBio lrRNA-seq from Genome-in-a-Bottle samples and identify methods that achieve high phasing accuracy and completeness. Our Snakemake pipeline leverages existing methods to perform variant calling, read-based phasing, transcript discovery, haplotype-resolved proteome construction, MS search, and downstream annotation. To demonstrate its utility, we apply the workflow to an induced pluripotent stem cell line (WTC11) and to an osteoblast differentiation time course. We show that sample-specific haplotype-resolved databases enable the detection of variant and splice peptides, allele-specific protein isoforms, and linked variants not detectable with reference proteomes. Together, our results demonstrate that lrRNA-seq-based phasing is feasible and effective for proteogenomics and provide a practical framework for allele-resolved proteome characterization in dynamic or disease-relevant settings.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Transcriptome-informed reduction of protein databases: an analysis of how and when proteogenomics enhances eukaryotic proteomics 94%
- Detecting haplotype-specific transcript variation in long reads with FLAIR2 93%
- Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics 93%
Similar papers in this journal
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 96%
- PROTRIDER: Protein abundance outlier detection from mass spectrometry-based proteomics data with a conditional autoencoder 95%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 94%
Similar papers in this journal
- The Integration of Proteogenomics and Ribosome Profiling Circumvents Key Limitations to Increase the Coverage and Confidence of Novel Microproteins 96%
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 95%
- Carafe enables high quality in silico spectral library generation for data-independent acquisition proteomics 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.