Splicing neoepitope prediction is sensitive to methodological differences
Prelot, L.; David, J. K.; Lin, A.; Kahles, A.; Yurchikova, M.; Thompson, R. F.; Rätsch, G.
Show abstract
MotivationCancer-specific neoepitopes may arise from abnormal splicing in the transcriptomic landscape (alternative splicing neoepitopes, ASNs), leading to divergent proteins with high potential immunogenicity. Predicting candidate ASNs requires a series of algorithmic and parameter choices, but no previous study has investigated the consistency and interpretability of these choices. ResultsWe apply two ASN prediction methods with 35 matched parameter sets to generate candidate ASNs for five breast and ovarian cancer samples from The Cancer Genome Atlas (TCGA). We find that: 1) junction 9-mer peptides generated by two similarly designed pipelines differed by, on average, 68.9% and 76.6% in the BRCA and OV cohorts, respectively, and putatively cancer-specific junction 9-mers from the two pipelines diverged further, by an average of 81.5 % in OV and 84.6 % in BRCA; 2) the most lenient filters in the BRCA cohort show the highest divergence, at 97 %; 3) the rate of mass spectrometry validation of ASNs protein presence in cells is dominated by the size of the input space; and 4) putatively cancer-specific ASNs found by the intersection of both pipelines can be validated when accounting for false discovery with an average of 1.74 and 284.4 candidates in the BRCA and OV cohorts, respectively. Taken together, these results highlight that ASN identification is fragile and difficult to reproduce across analysis platforms, with limited cross-pipeline overlap and strong dependence on parameter choices. Availability and ImplementationPython research code and scripts are available for download at https://github.com/ratschlab/projects2020_ohsu. Contactthompsre@ohsu.edu, raetsch@inf.ethz.ch
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 95%
- VIQoR: a web service for Visually supervised protein Inference and protein Quantification 95%
- Mistle: bringing spectral library predictions to metaproteomics with an efficient search index 95%
Similar papers in this journal
- The Personalized Proteome: Comparing Proteogenomics and Open Variant Search Approaches for Single Amino Acid Variant Detection 95%
- Massive proteogenomic reanalysis of publicly available proteomic datasets of human tissues in search for protein recoding via adenosine-to-inosine RNA editing 95%
- mokapot: Fast and flexible semi-supervised learning for peptide detection 95%
Similar papers in this journal
- Transcriptome-informed reduction of protein databases: an analysis of how and when proteogenomics enhances eukaryotic proteomics 94%
- RADAR: Differential analysis of MeRIP-seq data with a random effect model 92%
- DropletQC: improved identification of empty droplets and damaged cells in single-cell RNA-seq data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.