Back

Define protein variant functions with high-complexity mutagenesis libraries and enhanced mutation detection software ASMv1.0

Yang, X.; Hong, A. L.; Sharpe, T.; Giacomelli, A. O.; Lintner, R. E.; Alan, D.; Green, T.; Hayes, T. K.; Piccioni, F.; Fritchman, B.; Kawabe, H.; Sawyer, E.; Sprenkle, L.; Lee, B. P.; Persky, N. S.; Brown, A.; Greulich, H.; Aguirre, A. J.; Meyerson, M.; Hahn, W. C.; Johannessen, C. M.; Root, D. E.

2021-06-21 systems biology
10.1101/2021.06.16.448102 bioRxiv
Show abstract

Pooled variant expression libraries can test the phenotypes of thousands of variants of a gene in a single multiplexed experiment. In a library encoding all single-amino-acid substitutions of a protein, each variant differs from its reference only at a single codon-position located anywhere along the coding sequence. Consequently, accurately identifying these variants by sequencing is a major technical challenge. A popular but expensive brute-force approach is to divide the pool of variants into multiple smaller sub-libraries that each contains variants of a small region and that must each be constructed and screened individually, but that can then be PCR-amplified and fully sequenced with a single read to allow direct readout of variant abundance. Here we present an approach to screen very large variant libraries with mutations spanning a wide region in a single pool, including library design criteria and mutant-detection algorithms that permit reliable calling and counting of variants from large-scale sequencing data.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.