Mapping the Human Ghost Proteome: Classification and Experimental Detection Biases in the Identification of Alternative Microproteins
Montero-Calle, A.; Pelaez-Garcia, A.; Martin-Galiano, A. J.; Barderas, R.
Show abstract
The discovery of alternative proteins (AltProts), translated from non-canonical ORFs, has expanded the human proteome and revealed a hidden layer known as the "ghost proteome". Despite increasing evidence, AltProts detection remains challenging due to their small size, physicochemical heterogeneity, and lack of annotation. Here, we developed an integrated bioinformatic and proteomic workflow to benchmark the detection of reference proteins (RefProts), isoforms, and alternative microproteins (MicroAltProts) in colorectal cancer cells using four extraction protocols--HCl, RIPA buffer, RIPA with chloroform, and RIPA followed by 30 kDa filtration--combined with high-resolution data-independent acquisition mass spectrometry. We identified and quantified using the Orbitrap Astral mass spectrometer a total of 66,438 peptides corresponding to 12,584 different protein groups across methods, with RIPA-based extraction approaches providing the most comprehensive coverage. To reduce redundancy in the OpenProt database and focus on MicroAltProts, we curated the dataset by removing known isoforms and long proteins, yielding a non-redundant set of 183,937 MicroAltProts. K-means clustering based on eight ProtParam-derived features grouped MicroAltProts into four physicochemical clusters. Among them, 43 MicroAltProts (<200 amino acids) were experimentally validated by mass spectrometry and classified into tiers following recent recommended international guidelines. Cluster assignment of detected MicroAltProts revealed that HCl extraction favored disordered, alkaline proteins, while RIPA-based protocols enabled the identification of membrane-associated and amphipathic -helical MicroAltProts. Structural prediction indicated the presence of diverse folding determinants, including transmembrane helices, disordered regions, and nucleic acid-binding-like motifs. Altogether, this study provides a roadmap framework for the unbiased simultaneous detection of RefProts, isoforms, and AltProts, and supports a broader functional role for MicroAltProts.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 95%
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 95%
- An economic and robust TMT labeling approach for high throughput proteomic and metaproteomic analysis 94%
Similar papers in this journal
- 2-Mercaptoethanol/DMSO workflow enables highly reproducible quantitative proteomics 95%
- The E. coli PeptideAtlas Build: Characterizing the observed Escherichia coli pan-proteome and its post-translational modifications 95%
- Extensive and accurate benchmarking of DIA acquisition methods and software tools using a complex proteomic standard 95%
Similar papers in this journal
- vPro-MS enables identification of human-pathogenic viruses from patient samples by untargeted proteomics 95%
- Cross-platform Clinical Proteomics using the Charite Open Standard for Plasma Proteomics (OSPP) 95%
- High-throughput chemical proteomics workflow for profiling protein citrullination dynamics 95%
Similar papers in this journal
- A Benchmarking Framework for Comparative Evaluation of Low-Complexity Region Detection Tools in the Human Proteome 96%
- A Proteomic Platform to Identify Off-Target Proteins Associated with Therapeutic Modalities that Induce Protein Degradation or Gene Silencing 94%
- Low-resolution FAIMS for increased peptide coverage in low-load and single-cell proteomics 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.