Back

Variability in drought gene expression datasets highlight the need for community standardization

VanBuren, R.; Nguyen, A.; Marks, R. A.; Mercado, C.; Pardo, A.; Pardo, J.; Schuster, J.; St. Aubin, B.; Wilson, M. L.; Rhee, S. Y.

2024-02-06 plant biology
10.1101/2024.02.04.578814 bioRxiv
Show abstract

Physiologically relevant drought stress is difficult to apply consistently, and the heterogeneity in experimental design, growth conditions, and sampling schemes make it challenging to compare water deficit studies in plants. Here, we re-analyzed hundreds of drought gene expression experiments across diverse model and crop species and quantified the variability across studies. We found that drought studies are surprisingly uncomparable, even when accounting for differences in genotype, environment, drought severity, and method of drying. Many studies, including most Arabidopsis work, lack high-quality phenotypic and physiological datasets to accompany gene expression, making it impossible to assess the severity or in some cases the occurrence of water deficit stress events. From these datasets, we developed supervised learning classifiers that can accurately predict if RNA-seq samples have experienced a physiologically relevant drought stress, and suggest this can be used as a quality control for future studies. Together, our analyses highlight the need for more community standardization, and the importance of paired physiology data to quantify stress severity for reproducibility and future data analyses.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.