Fungal Pathogen Gene Selection for Predicting the Onset of Infection Using a Multi-Stage Machine Learning Approach
Thomas, G.; Stoner, O.; Costa, F.; Ames, R. M.
Show abstract
Phytopathogenic fungi pose a serious threat to global food security. Next-generation sequencing technologies, such as transcriptomics, are increasingly used to profile infection, assess environmental adaptation and gauge host-responses. The accumulation of these large-scale data has created the opportunity to employ new computational methods to gain greater biological insights. Machine learning approaches, that learn to identify patterns in complex data sets, have recently been applied to the field of plant-pathogen interactions. Here, we apply a machine learning approach to transcriptomics data for the fungal pathogen Zymoseptoria tritici, to predict the onset of infection as measured by timing of the appearance of necrosis. We present a method for identifying the most important genes that predict infection timings, accurately classify isolates as early and late infectors and predict the timing of infection of novel isolates using only a subset of the data. These methods and the genes identified further demonstrate the use of these tools in the field of plant-pathogen interactions and have implications for the identification of biomarkers for disease monitoring and forecasting. Fungi that infect plants pose a serious threat to global food security. Methods to study these pathogens generate vast amounts of data that create new opportunities for computational tools to analyse them. Machine learning methods can learn patterns in complex data such as when genes are turned on or off in fungal plant pathogens. In this study we use machine learning approaches to predict the onset of infection in several isolates of an important fungal pathogen. We show that these methods can identify a small group of genes that are predictive of the infection onset. We can even use these methods on novel isolates to infer the likely timing of disease development. Our work has implications for plant disease diagnosis, monitoring and forecasting.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genome reconstruction of the non-culturable spinach downy mildew Peronospora effusa by metagenome filtering 95%
- Multi-locus phylogenetic network analysis of Ampelomyces mycoparasites isolated from diverse powdery mildews in Australia and the generation of two de novo genome assemblies 93%
- Whole genome resequencing and comparative genome analysis of three Puccinia striiformis f. sp. tritici pathotypes prevalent in India 93%
Similar papers in this journal
- Comparative genomics of Alternaria species provides insights into the pathogenic lifestyle of Alternaria brassicae - a pathogen of the Brassicaceae family 93%
- Comparative genomic analyses shed light on the introduction routes of rice-pathogenic Burkholderia gladioli strains into Bangladesh 93%
- Analyses of Xenorhabdus griffiniae genomes reveal two distinct sub-species that display intra-species variation due to prophages. 92%
Similar papers in this journal
- Effectors with different gears: divergence of Ustilago maydis effector genes is associated with their temporal expression pattern during plant infection 94%
- The NADPH Oxidase A of Verticillium dahliae is Essential for Pathogenicity, Normal Development, and Stress Tolerance, and it Interacts with Yap1 to Regulate Redox Homeostasis 93%
- The Ustilago hordei-barley interaction is a versatile system to characterize fungal effectors 93%
Similar papers in this journal
- The Venturia inaequalis effector repertoire is expressed in waves and is dominated by expanded families with predicted structural similarity to avirulence proteins from other plant-pathogenic fungi 95%
- Large-scale transcriptomics to dissect two years of the life of a fungal phytopathogen interacting with its host plant 95%
- Uncovering the history of recombination and population structure in western Canadian stripe rust populations through mating-type alleles 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.