Genome-wide predictions of genetic redundancy in Arabidopsis thaliana
Cusack, S. A.; Wang, P.; Moore, B. M.; Meng, F.; Conner, J. K.; Krysan, P. J.; Lehti-Shiu, M. D.; Shiu, S.-H.
Show abstract
Genetic redundancy refers to a situation where an individual with a loss-of-function mutation in one gene (single mutant) does not show an apparent phenotype until one or more paralogs are also knocked out (double/higher-order mutant). Previous studies have identified some characteristics common among redundant gene pairs, but a predictive model of genetic redundancy incorporating a wide variety of features has not yet been established. In addition, the relative importance of these characteristics for genetic redundancy remains unclear. Here, we establish machine learning models for predicting whether a gene pair is likely redundant or not in the model plant Arabidopsis thaliana. Benchmark gene pairs were classified based on six feature categories: functional annotations, evolutionary conservation including duplication patterns and mechanisms, epigenetic marks, protein properties including post-translational modifications, gene expression, and gene network properties. The definition of redundancy, data transformations, feature subsets, and machine learning algorithms used affected model performance significantly. Among the most important features in predicting gene pairs as redundant were having a paralog(s) from recent duplication events, annotation as a transcription factor, downregulation during stress conditions, and having similar expression patterns under stress conditions. Predictions were then tested using phenotype data withheld from model building and validated using well-characterized, redundant and nonredundant gene pairs. This genetic redundancy model sheds light on characteristics that may contribute to long-term maintenance of paralogs that are seemingly functionally redundant, and will ultimately allow for more targeted generation of functionally informative double mutants, advancing functional genomic studies.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Species-specific gene duplication in Arabidopsis thaliana evolved novel phenotypic effects on morphological traits under strong positive selection 95%
- Systematic histone H4 replacement in Arabidopsis thaliana reveals a role for H4R17 in regulating flowering time 94%
- Single-plant-omics reveals the cascade of transcriptional changes during the vegetative-to-reproductive transition 94%
Similar papers in this journal
- A Decoy Library Uncovers U-box E3 Ubiquitin Ligases that Regulate Flowering Time in Arabidopsis 94%
- Genome-wide association mapping of transcriptome variation in Mimulus guttatus indicates differing patterns of selection on cis- versus trans-acting mutations 93%
- Base-pairing requirements for small RNA-mediated gene silencing of recessive self-incompatibility alleles in Arabidopsis halleri. 93%
Similar papers in this journal
Similar papers in this journal
- Genome-wide misexpression associated with hybrid sterility in Mimulus (monkeyflower) 94%
- Whole-genome duplications and the long-term evolution of gene regulatory networks in angiosperms 94%
- The evolutionary forces shaping cis and trans regulation of gene expression within a population of outcrossing plants. 94%
Similar papers in this journal
- Optimizing the use of gene expression data to predict plant metabolic pathway memberships 95%
- A likelihood ratio test for detecting shifts in homeolog expression ratios in allopolyploids 94%
- Revisiting regulatory decoherence and phenotypic integration: accounting for temporal bias in co-expression analyses 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.