Back

Resistify - A rapid and accurate annotation tool to identify NLRs and study their genomic organisation

Smith, M.; Jones, J. T.; Hein, I.

2024-02-16 plant biology
10.1101/2024.02.14.580321 bioRxiv
Show abstract

BackgroundNucleotide-binding domain Leucine-rich Repeat (NLR) proteins are a key component of the plant innate immune system. In plant genomes NLRs exhibit considerable presence/absence variation and sequence diversity. Recent advances in sequencing technologies have made the generation of high-quality novel plant genome assemblies considerably more straightforward. Accurately identifying NLRs from these genomes is a prerequisite for improving our understanding of NLRs and identifying potential novel sources of disease resistance. ResultsWhilst several tools have been developed to predict NLRs, they are hampered by low accuracy, speed, and availability. Here, the NLR annotation tool Resistify is presented. Resistify is an easy-to-use, rapid, and accurate tool to identify and classify NLRs from protein sequences. Applying Resistify to the RefPlantNLR database demonstrates that it can correctly identify NLRs from a diverse range of species. Applying Resistify in combination with tools to identify transposable elements to a panel of Solanaceae genomes reveals a previously undescribed association between NLRs and Helitron transposable elements. ConclusionResistify can rapidly identify NLRs within plant genomes and provides accurate structural classifications. Its ease of use and accessibility allows easy integration into bioinformatic workflows and projects, enhancing the study of this important group of genes. Applying Resistify to a Solanaceae pangenome reveals an undescribed association between NLRs and transposable elements. Availability: https://github.com/SwiftSeal/resistify

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.