P3ANUT: An enhanced DNA sequencing analysis platform for uncovering and correcting errors in peptide phage display library screening
Tucker, L.; Kolland, E.; Gooding, J.; Boudjelal, H.; Ahmadipour, S.; Field, R. A.; Warren, D. T.; Baker, D.; Ma, Y.; Marin, M. J.; Wu, T.; Morris, C. J.
Show abstract
Display technologies are used extensively in the discovery of peptides and antibodies towards the development of new medicines and diagnostic tools. Phage display technology enables the filtering of <1010 random peptides/antibodies down to and enriched pool of candidates with favourable binding affinities. In recent years, next-generation DNA sequencing technologies have increased the precision and accuracy with which the peptide sequences of phage display library clones are identified. Inaccuracies in DNA sequencing such as substitutions, insertions and deletions in the library oligonucleotide region have the potential to result in the identification of erroneous candidate sequences. Here, we describe a Python Pipeline for Phage Analysis through a Normative Unified Toolset (P3ANUT) which employs Levenshtein distance, k-mer approaches and a novel encoding scheme on paired-end sequencing outputs to correct sequencing errors from next-generation sequencing outputs of display library screens. We introduce an easy-to-use and highly customisable computational tool with graphical user- and command line interfaces to process entire datasets within a single input, as well as visualisation tools for candidate analysis and data generation. P3ANUT shows significant improvements in read recovery, overall read quality, and runtime compared to a previously published pipeline. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=96 SRC="FIGDIR/small/648809v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@15e1172org.highwire.dtl.DTLVardef@cb9097org.highwire.dtl.DTLVardef@81cabcorg.highwire.dtl.DTLVardef@1253865_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TemStaPro: protein thermostability prediction using sequence representations from protein language models 94%
- Variant Library Annotation Tool (VaLiAnT): an oligonucleotide library design and annotation tool for Saturation Genome Editing and other Deep Mutational Scanning experiments 94%
- Rapid T cell receptor interaction grouping with ting 94%
Similar papers in this journal
Similar papers in this journal
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 94%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 94%
- A generalised protein identification method for novel and diverse sequencing technologies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.