Back

AliceDB database and pipeline for identification of natural protein variants based on mass spectrometry measurement data

Thiel, M.; Rozycka, A.; Puchalski, M.; Oldziej, S.

2026-06-15 bioinformatics
10.64898/2026.06.11.731579 bioRxiv
Show abstract

The natural variation that distinguishes living organisms within a single species is currently being studied intensively, primarily at the genetic level. Unfortunately, studies of natural variants at the level of protein gene products are not very common, mainly due to the lack of appropriate databases and bioinformatics tools. The main research technique used to study proteomes/peptidomes is mass spectrometry (MS). A classic method for interpreting raw mass spectrometry data in proteomic/peptidomic studies involves the use of databases containing representative (canonical) sequences that define the proteome of the organism under study. In this paper, we present the AliceDB database, which contains information on over 7 million natural variants of protein sequences described in the scientific literature for Homo sapiens. The data contained in the AliceDB database can be utilized using widely available and commonly used software for interpreting proteomic data. Test results regarding the use of the AliceDB database for the interpretation of proteomic data indicate that accounting for the presence of natural variants increases both the number and quality of identified proteins. Furthermore, it is easy to identify protein sequence variants that may, for example, be of significance in medicine.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Journal of Proteome Research
234 papers in training set
Top 0.1%
28.1%
2
PROTEOMICS
43 papers in training set
Top 0.1%
13.6%
3
Molecular & Cellular Proteomics
25 papers in training set
Top 0.1%
8.3%
50% of probability mass above
4
Journal of Proteomics
28 papers in training set
Top 0.1%
8.3%
5
PLOS ONE
5266 papers in training set
Top 27%
5.8%
6
BMC Bioinformatics
457 papers in training set
Top 3%
2.5%
7
Journal of the American Society for Mass Spectrometry
37 papers in training set
Top 0.2%
2.2%
8
Bioinformatics
1204 papers in training set
Top 6%
2.0%
9
Analytical Chemistry
218 papers in training set
Top 1%
1.8%
10
Biochimica et Biophysica Acta (BBA) - Biomembranes
36 papers in training set
Top 0.2%
1.6%
11
Scientific Reports
3612 papers in training set
Top 58%
1.5%
12
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.2%
13
Frontiers in Bioinformatics
49 papers in training set
Top 0.7%
1.2%
14
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.2%
15
PeerJ
308 papers in training set
Top 9%
1.1%
16
Genes
144 papers in training set
Top 4%
0.9%
17
Molecular & Cellular Proteomics
158 papers in training set
Top 1%
0.6%
18
International Journal of Molecular Sciences
494 papers in training set
Top 16%
0.6%
19
ACS Omega
105 papers in training set
Top 4%
0.6%
20
Scientific Data
209 papers in training set
Top 4%
0.5%
21
SoftwareX
15 papers in training set
Top 0.4%
0.5%
22
Analytical and Bioanalytical Chemistry
18 papers in training set
Top 0.4%
0.5%
23
International Journal of Biological Macromolecules
76 papers in training set
Top 2%
0.5%
24
Molecules
39 papers in training set
Top 2%
0.5%
25
GigaScience
212 papers in training set
Top 6%
0.5%
26
BMC Genomics
406 papers in training set
Top 10%
0.5%
27
PLOS Computational Biology
1863 papers in training set
Top 23%
0.5%