Back

Sample-specific haplotype-resolved protein isoform characterization via long-read RNA-seq-based proteogenomics

Wissel, D.; Sheynkman, G. M.; Robinson, M. D.

2026-03-04 bioinformatics
10.1101/2025.11.21.689818 bioRxiv
Show abstract

Protein isoform inference from bottom-up mass spectrometry (MS) relies on database search strategies that assume the reference protein database accurately reflects the full repertoire of genetic and transcriptomic states present in the sample being analyzed. Long-read RNA sequencing (lrRNA-seq) now enables simultaneous recovery of complete transcript (splice) structures and the genetic variants present on each molecule, offering a direct route to allele-specific isoforms. Yet, this capability has not been fully leveraged to improve MS-based proteogenomics workflows. Here, we develop an end-to-end workflow for constructing and searching haplotype-resolved, sample-specific proteomes using matched lrRNA-seq and MS data. We benchmark phasing algorithms on PacBio lrRNA-seq from Genome-in-a-Bottle samples and identify methods that achieve high phasing accuracy and completeness. Our Snakemake pipeline leverages existing methods to perform variant calling, read-based phasing, transcript discovery, haplotype-resolved proteome construction, MS search, and downstream annotation. To demonstrate its utility, we apply the workflow to an induced pluripotent stem cell line (WTC11) and to an osteoblast differentiation time course. We show that sample-specific haplotype-resolved databases enable the detection of variant and splice peptides, allele-specific protein isoforms, and linked variants not detectable with reference proteomes. Together, our results demonstrate that lrRNA-seq-based phasing is feasible and effective for proteogenomics and provide a practical framework for allele-resolved proteome characterization in dynamic or disease-relevant settings.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.