Back

P3ANUT: An enhanced DNA sequencing analysis platform for uncovering and correcting errors in peptide phage display library screening

Tucker, L.; Kolland, E.; Gooding, J.; Boudjelal, H.; Ahmadipour, S.; Field, R. A.; Warren, D. T.; Baker, D.; Ma, Y.; Marin, M. J.; Wu, T.; Morris, C. J.

2025-05-12 bioinformatics
10.1101/2025.05.12.648809 bioRxiv
Show abstract

Display technologies are used extensively in the discovery of peptides and antibodies towards the development of new medicines and diagnostic tools. Phage display technology enables the filtering of <1010 random peptides/antibodies down to and enriched pool of candidates with favourable binding affinities. In recent years, next-generation DNA sequencing technologies have increased the precision and accuracy with which the peptide sequences of phage display library clones are identified. Inaccuracies in DNA sequencing such as substitutions, insertions and deletions in the library oligonucleotide region have the potential to result in the identification of erroneous candidate sequences. Here, we describe a Python Pipeline for Phage Analysis through a Normative Unified Toolset (P3ANUT) which employs Levenshtein distance, k-mer approaches and a novel encoding scheme on paired-end sequencing outputs to correct sequencing errors from next-generation sequencing outputs of display library screens. We introduce an easy-to-use and highly customisable computational tool with graphical user- and command line interfaces to process entire datasets within a single input, as well as visualisation tools for candidate analysis and data generation. P3ANUT shows significant improvements in read recovery, overall read quality, and runtime compared to a previously published pipeline. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=96 SRC="FIGDIR/small/648809v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@15e1172org.highwire.dtl.DTLVardef@cb9097org.highwire.dtl.DTLVardef@81cabcorg.highwire.dtl.DTLVardef@1253865_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.