Back

moPepGen: Rapid and Comprehensive Proteoform Identification

Zhu, C.; Liu, L. Y.; Yamaguchi, T. N.; Zhu, H.; Hugh-White, R.; Livingstone, J.; Patel, Y.; Kislinger, T.; Boutros, P. C.

2024-03-31 bioinformatics
10.1101/2024.03.28.587261 bioRxiv
Show abstract

Proteogenomics is limited by challenges of modeling the complexities of gene expression. We create moPepGen, a graph-based algorithm that comprehensively generates non-canonical peptides in linear time. moPepGen works with multiple technologies, in multiple species and on all types of genetic and transcriptomic data. In human cancer proteomes, it enumerates previously unobservable noncanonical peptides arising from germline and somatic genomic variants, noncoding open reading frames, RNA fusions and RNA circularization.

Published in Nature Biotechnology (predicted rank #4) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.