Back

Whole protein sequencing and quantification without proteolysis, terminal residue cleavage, or purification: A computational model

Sampath, G.

2024-03-14 bioengineering
10.1101/2024.03.13.584825 bioRxiv
Show abstract

Sequencing and quantification of whole proteins in a sample without separation, terminal residue cleavage, or proteolysis are modeled computationally. Similar to recent work on DNA sequencing (PNAS 113, 5233-5238, 2016), a high-volume conjugate is attached to every instance of amino acid (AA) type AAi, 1 [≤] i [≤] 20, in an unfolded whole protein, which is then translocated through a nanopore. From the volume excluded by 2L residues in a pore of length L nm (a proxy for the blockade current), a partial sequence containing AAi is obtained. Translocation is assumed to be unidirectional, with residues exiting the pore at a roughly constant rate of [~]1/s (Nature Biotechnology 41, 1130-1139, 2023). The blockade signal is sampled at intervals of 1 s and digitized with a step precision of 70 nm3; the positions of the AAis are obtained from the positions of well-defined quantum jumps in the signal. This procedure is applied to all 20 standard AA types, the resulting 20 partial sequences are merged to obtain the whole protein sequence. The complexity of subsequence computation is O(N) for a protein with N residues. The method is illustrated with a sample protein from the human proteome (Uniprot id UP000005640_9606). A mixture of M protein molecules (including multiple copies) can be sequenced by constructing an M x 20 array of partial sequences from which proteins occurring multiple times are first isolated and their sequences obtained separately. The remaining M singly-occurring molecules are detected from M disjoint paths through the 20 columns of the reduced M x 20 array. Detection complexity is O(M20), which is nominally in polynomial time but practical only for small M; to use this method a sample may be subdivided into subsamples down to this level. Quantification of proteins can be done by sorting their computed sequences on the sequence strings and counting the number of duplicates. The possibility of translating this procedure into practice and related implementation issues are discussed.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 8%
19.3%
2
Methods
34 papers in training set
Top 0.1%
13.5%
3
Bioinformatics
1204 papers in training set
Top 2%
12.4%
4
Scientific Reports
3612 papers in training set
Top 12%
6.5%
50% of probability mass above
5
International Journal of Molecular Sciences
494 papers in training set
Top 3%
3.4%
6
PLOS Computational Biology
1863 papers in training set
Top 9%
3.4%
7
Frontiers in Molecular Biosciences
102 papers in training set
Top 0.3%
2.9%
8
Journal of Open Source Software
25 papers in training set
Top 0.1%
2.8%
9
Royal Society Open Science
214 papers in training set
Top 2%
2.2%
10
BMC Genomics
406 papers in training set
Top 3%
2.2%
11
Biophysical Journal
631 papers in training set
Top 3%
2.0%
12
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 0.5%
1.6%
13
MethodsX
16 papers in training set
Top 0.1%
1.4%
14
Journal of The Royal Society Interface
235 papers in training set
Top 3%
1.4%
15
F1000Research
88 papers in training set
Top 2%
1.2%
16
PeerJ
308 papers in training set
Top 7%
1.2%
17
Synthetic Biology
24 papers in training set
Top 0.2%
1.1%
18
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 2%
0.9%
19
SoftwareX
15 papers in training set
Top 0.2%
0.9%
20
IFAC-PapersOnLine
13 papers in training set
Top 0.2%
0.6%
21
International Journal for Numerical Methods in Biomedical Engineering
14 papers in training set
Top 0.3%
0.6%
22
Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences
12 papers in training set
Top 0.2%
0.5%
23
ACS Synthetic Biology
287 papers in training set
Top 3%
0.5%
24
Mathematical Biosciences
49 papers in training set
Top 1%
0.5%
25
Frontiers in Pharmacology
111 papers in training set
Top 4%
0.5%