Back

Identification of Diverse Cellulose Binding Domains using in silico Prioritisation and High-Throughput Screening

Field, C. R.; Mattey, A. P.; Cosgrove, S. C.; Hay, S.

2025-10-07 bioinformatics
10.1101/2025.10.06.677802 bioRxiv
Show abstract

In silico pre-screening methods are becoming an increasingly important step in making large scale enzyme screening experiments manageable, particularly in the sampling large protein sequence datasets for maximum diversity. Here, we develop a bioinformatics workflow that utilises automated gene sequence annotation, substrate prediction and structure prediction tools to effectively isolate representative protein sequences, thereby maximising the information gathered on the wider dataset from relatively few in vitro experiments. Included in this workflow is a new tool, ColabAlign, which performs pairwise structural alignments to construct a structure-informed dendrogram, from which representatives are selected based on clustering. The workflow was applied to the identification of cellulose-binding carbohydrate-binding modules (CBMs) suitable for enzyme immobilisation tag development. 47 sequentially and structurally diverse mCherry-CBM fusions were tested for binding against cellulose, chitin and spent coffee grounds (SCG) using a pulldown assay from cell lysate. We successfully identified 5 CBMs with significant binding towards commercial and waste cellulose support materials, suitable for further tag design work. This work also provides the first experimental evidence of chitin binding by one or more members of CBM6 and CBM46.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.