Back

A deep audit of the PeptideAtlas database uncovers evidence for unannotated coding genes and aberrant translation

Rodriguez, J. M.; Maquedano, M.; Cerdan-Velez, D.; Calvo, E.; Vazquez, J.; Tress, M. L.

2024-11-15 genomics
10.1101/2024.11.14.623419 bioRxiv
Show abstract

The human genome has been the subject of intense scrutiny by experimental and manual curation projects for more than two decades. Novel coding genes have been proposed from large-scale RNASeq, ribosome profiling and proteomics experiments. Here we carry out an in-depth analysis of an entire proteomics database. We analysed the proteins, peptides and spectra housed in the human build of the PeptideAtlas proteomics database to identify coding regions that are not yet annotated in the GENCODE reference gene set. We find support for hundreds of missing alternative protein isoforms and unannotated upstream translations, and evidence of cross-contamination from other species. There was reliable peptide evidence for 34 novel unannotated open reading frames (ORFs) in PeptideAtlas. We find that almost half belong to coding genes that are missing from GENCODE and other reference sets. Most of the remaining ORFs were not conserved beyond human, however, and their peptide confirmation was restricted to cancer cell lines. We show that this is strong evidence for aberrant translation, raising important questions about the extent of aberrant translation and how these ORFs should be annotated in reference genomes.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
PROTEOMICS
43 papers in training set
Top 0.1%
19.1%
2
PLOS ONE
5266 papers in training set
Top 9%
19.1%
3
BMC Genomics
406 papers in training set
Top 0.4%
8.1%
4
PLOS Computational Biology
1863 papers in training set
Top 5%
6.9%
50% of probability mass above
5
Journal of Proteome Research
234 papers in training set
Top 0.6%
5.3%
6
Genome Biology
637 papers in training set
Top 3%
4.2%
7
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.0%
8
Genomics, Proteomics & Bioinformatics
172 papers in training set
Top 0.9%
2.0%
9
Scientific Reports
3612 papers in training set
Top 50%
2.0%
10
Mobile DNA
31 papers in training set
Top 0.1%
1.8%
11
International Journal of Molecular Sciences
494 papers in training set
Top 8%
1.7%
12
BMC Medical Genomics
50 papers in training set
Top 0.6%
1.5%
13
Methods
34 papers in training set
Top 0.3%
1.5%
14
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.5%
15
eLife
5828 papers in training set
Top 56%
1.2%
16
Molecular & Cellular Proteomics
158 papers in training set
Top 1%
1.2%
17
Genomics
64 papers in training set
Top 1%
1.2%
18
Life Science Alliance
285 papers in training set
Top 5%
1.1%
19
GigaScience
212 papers in training set
Top 4%
1.0%
20
Genes
144 papers in training set
Top 4%
0.9%
21
BMC Genomic Data
13 papers in training set
Top 0.1%
0.9%
22
Journal of Molecular Biology
232 papers in training set
Top 3%
0.9%
23
Nucleic Acids Research
1281 papers in training set
Top 14%
0.6%
24
Nature Communications
5641 papers in training set
Top 58%
0.6%