Back

Identifying residues in unfolded whole proteins with a nanopore: a theoretical model based on linear inequalities

Sampath, G.

2023-09-03 bioengineering
10.1101/2023.08.31.555759 bioRxiv
Show abstract

A theoretical model is proposed for the identification of individual amino acids (AAs) in an unfolded whole proteins primary sequence. It is based in part on a recent report (Nat. Biotech. 41, 1130-1139, 2023) that describes the unfolding and translocation of whole proteins at constant speed through a biological nanopore (alpha-Hemolysin) of length 5 nm with a residue dwell time inside the pore of [~]10 s. Here current blockade levels in the pore due to the translocating protein are assumed to be measured with a limited precision of 70 nm3 and a bandwidth of 20 KHz for measurement with a low-bandwidth detector. Exclusion volumes in two pores of slightly different lengths are used as a computational proxy for the blockade signal; subsequence exclusion volume differences along the protein sequence are computed from the sampled translocation signals in the two pores relatively shifted multiple times. These are then converted into a system of linear inequalities that can be solved with linear programming and related methods; residues are coarsely identified as belonging to one of 4 subsets of the 20 standard AAs. To obtain the exact identity of a residue an artifice analogous to the use of base-specific tags for DNA sequencing with a nanopore (PNAS 113, 5233-5238, 2016) is used. Conjugates that add volume are attached to a given AA type, this biases the set of inequalities toward the volume of the conjugated AA, from this biased set the position of occurrence of every residue of the AA type in the whole sequence is extracted. By applying this step separately to each of the 20 standard AAs the full sequence can be obtained. The procedure is illustrated with a protein in the human proteome (Uniprot id UP000005640_9606).

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 9%
19.1%
2
Methods
34 papers in training set
Top 0.1%
17.5%
3
Scientific Reports
3612 papers in training set
Top 9%
6.9%
4
International Journal of Molecular Sciences
494 papers in training set
Top 1%
5.6%
5
PLOS Computational Biology
1863 papers in training set
Top 7%
5.0%
50% of probability mass above
6
Biophysical Journal
631 papers in training set
Top 2%
2.5%
7
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 0.4%
1.8%
8
Frontiers in Molecular Biosciences
102 papers in training set
Top 0.6%
1.8%
9
Analytical and Bioanalytical Chemistry
18 papers in training set
Top 0.2%
1.7%
10
Royal Society Open Science
214 papers in training set
Top 3%
1.7%
11
Bioinformatics
1204 papers in training set
Top 7%
1.5%
12
MethodsX
16 papers in training set
Top 0.1%
1.2%
13
SoftwareX
15 papers in training set
Top 0.2%
1.2%
14
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 2%
1.1%
15
Sensors
43 papers in training set
Top 1%
1.0%
16
International Journal for Numerical Methods in Biomedical Engineering
14 papers in training set
Top 0.2%
0.9%
17
Analytical Biochemistry
26 papers in training set
Top 0.4%
0.9%
18
Synthetic Biology
24 papers in training set
Top 0.3%
0.9%
19
Journal of The Royal Society Interface
235 papers in training set
Top 4%
0.9%
20
Biochimica et Biophysica Acta (BBA) - Biomembranes
36 papers in training set
Top 0.3%
0.9%
21
Mathematical Biosciences and Engineering
23 papers in training set
Top 0.9%
0.6%
22
BMC Genomics
406 papers in training set
Top 9%
0.6%
23
Bioinformatics Advances
203 papers in training set
Top 5%
0.6%
24
Biology
45 papers in training set
Top 1%
0.6%
25
Journal of Open Source Software
25 papers in training set
Top 0.4%
0.6%
26
Journal of Proteome Research
234 papers in training set
Top 2%
0.6%
27
Physical Biology
46 papers in training set
Top 1%
0.6%
28
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%
29
Communications Chemistry
48 papers in training set
Top 2%
0.6%
30
Microbiology Spectrum
469 papers in training set
Top 11%
0.6%