XMAn Update - A Database of Homo sapiens Mutated Peptides
Haueis, J. R. S.; Lazar, I. M.
Show abstract
Mass spectrometry (MS) is the leading technology for identifying proteins in complex biological samples. It relies on the use of tandem MS alongside a reference database of canonical protein sequences to computationally identify peptides and their parent proteins. The canonical sequences represent the most widely expressed and functionally validated forms of proteins. Consequently, disease-induced or disease-supportive variants, such as those associated with cancer, will evade detection if they are absent from the database. To address this challenge, this study introduces a revised release of the Unkown Mutation Analysis (XMAn) database by incorporating coding missense and nonsense mutations from the latest versions (v103) of the COSMIC Genome Screen Mutants (GSM) and Cancer Gene Census (CGC) datasets in two distinct FASTA-formatted peptide databases comprising 3,848,499 and 312,658 variants, respectively. The mutated peptides were matched to reviewed, non-redundant UniProt Homo sapiens protein entries (18,362 and 746), and characterized in terms of nucleotide- and amino acid mutation frequencies, peptide length distributions, and associations between specific single-nucleotide (SNV) and single amino acid (SAAVs) variants. Applied to the analysis of MDA-MB-231 breast cancer cell-membrane protein fractions, the database enabled the identification of 300+ high-quality variant peptides - several localized to functional protein-binding and catalytic domains - and 23 aberrant protein products mapped to the CGC dataset. The database is hosted and available for download on Zenodo (XMAn/gsm doi: 10.5281/zenodo.21781023; XMAn/cgc doi: 10.5281/zenodo.21781514) or can be accessed through https://sites.google.com/vt.edu/xman-db/home.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Fast and memory efficient searching of large-scale mass spectrometry data using Tide 95%
- PAMPA: a software for peptide markers and taxonomic identification for ZooMS samples in Archaeology and Paleontology 95%
- Extremely fast and accurate open modification spectral library searching of high-resolution mass spectra using feature hashing and graphics processing units 95%
Similar papers in this journal
- Improved peptide search for identification of SUMO and sequence-based modifications, in MaxSBM 97%
- MS-EmpiRe utilizes peptide-level noise distributions for ultra sensitive detection of differentially abundant proteins 94%
- MaXLinker: proteome-wide cross-link identifications with high specificity and sensitivity 94%
Similar papers in this journal
- DrawAlignR: An interactive tool for across run chromatogram alignment visualization 94%
- MMS2plot: an R package for visualizing multiple MS/MS spectra for groups of modified and non-modified peptides 94%
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 94%
Similar papers in this journal
- HXMS: a standardized file format for HX/MS data 95%
- Generation of ENSEMBL-based proteogenomics databases boosts the identification of non-canonical peptides. 95%
- Alpha-XIC: a deep neural network for scoring the coelution of peak groups improves peptide identification by data-independent acquisition mass spectrometry 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.