Back

A minimalist binary/digital approach to large-scale single molecule protein identification with optically labeled tRNAs and multiple carboxypeptidases and its extension to peptide sequencing

Sampath, G.

2024-12-05 bioengineering
10.1101/2024.12.02.626402 bioRxiv
Show abstract

Recently a binary/digital scheme based on the superspecificity property of transfer RNAs (tRNAs) was proposed for the identification of single amino acids (AAs) from binary-valued measurements (Eur. Phys. J. E 45, 94, 2022). There are two formulations, they can be used to sequence short peptides and/or identify their parent proteins. In one of them an array of peptides is sequenced in 20 cycles by adding 20 different tRNAs carrying a fluorescent tag, optically recognizing the C-terminal residues, and cleaving the latter with a carboxypeptidase; the process is repeated over the peptides in parallel. Here this scheme is used to develop in theory a minimalist approach to protein identification that uses only two tRNAs and the carboxypeptidases A, B, and C. The latter form a complete and mutually exclusive set capable of cleaving all 20 AA types; this divides the 20 AAs into three classes. The sequences obtained are partial sequences in the reduced alphabet, their parent proteins can be obtained by search through a proteome database. The AA class of the terminal residue of every peptide in the array can be identified in a single cycle by using the three carboxypeptidases in the order C-B-A. With peptide lengths of [~]20 and a cycle time of [~]1 hour, the parent proteins of K peptides can be obtained in about 20 hours. This is independent of K (within the limits imposed by the imaging method used) and the dynamic range of a proteome; thus in theory a whole proteome can be processed in less than a day. Computational results suggest that the parent proteins of over 92% of peptides from the human proteome (Uniprot id UP000005640_9606) can be identified. The identification rate when residues are skipped due to carboxypeptidases cleaving the second and later residues in delayed reactions is about [~]90% with 1 or 2 skips. Full sequencing without skipped residues can be done by using all 20 tRNA types over 20 cycles in increasing order of cleavage time of the 20 AA types; a recursive procedure is given.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Methods
34 papers in training set
Top 0.1%
20.1%
2
PLOS ONE
5266 papers in training set
Top 8%
20.1%
3
Journal of Proteome Research
234 papers in training set
Top 0.5%
7.3%
4
International Journal of Molecular Sciences
494 papers in training set
Top 0.7%
6.8%
50% of probability mass above
5
Scientific Reports
3612 papers in training set
Top 21%
4.7%
6
PLOS Computational Biology
1863 papers in training set
Top 11%
2.7%
7
BMC Genomics
406 papers in training set
Top 3%
2.3%
8
Analytical and Bioanalytical Chemistry
18 papers in training set
Top 0.1%
1.8%
9
Frontiers in Molecular Biosciences
102 papers in training set
Top 0.8%
1.6%
10
Bioinformatics
1204 papers in training set
Top 7%
1.6%
11
Journal of the American Society for Mass Spectrometry
37 papers in training set
Top 0.4%
1.2%
12
New Biotechnology
12 papers in training set
Top 0.1%
1.2%
13
SoftwareX
15 papers in training set
Top 0.2%
1.2%
14
Royal Society Open Science
214 papers in training set
Top 4%
1.2%
15
Analytical Chemistry
218 papers in training set
Top 2%
0.9%
16
Sensors
43 papers in training set
Top 1%
0.9%
17
PeerJ
308 papers in training set
Top 12%
0.7%
18
Frontiers in Bioinformatics
49 papers in training set
Top 2%
0.7%
19
Journal of Open Source Software
25 papers in training set
Top 0.4%
0.7%
20
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 3%
0.7%
21
PROTEOMICS
43 papers in training set
Top 1.0%
0.5%
22
BMC Bioinformatics
457 papers in training set
Top 6%
0.5%
23
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.5%
24
ACS Omega
105 papers in training set
Top 4%
0.5%
25
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 1%
0.5%
26
F1000Research
88 papers in training set
Top 5%
0.5%
27
Synthetic Biology
24 papers in training set
Top 0.4%
0.5%
28
Gigabyte
62 papers in training set
Top 1%
0.5%