Predicting plant traits using large conglomerate RNA-seq datasets
Hadish, J. A. A.; Honaas, L.; Ficklin, S.
Show abstract
RNA-seq datasets offer potential for predicting phenotypic traits, but the optimal dataset size for reliable predictions remains unclear. To explore this question, we compiled large-scale transcriptomic datasets across 12 plant species and used them to predict the phenotypic variables of tissue type and age. Predictions were made using random forest models. We created these models using an increasing number of samples to create performance curves. These curves show that only a few hundred samples are required to achieve maximum accuracy for tissue classification, while predicting age demands a few thousand samples. Acceptable prediction accuracy can be achieved at even lower numbers of samples. Our findings provide a benchmark for designing transcriptomic studies aimed at phenotype prediction and highlight the differing complexities involved in predicting simple versus more complex traits. CORE IDEAS- Plant transcriptomes are highly dynamic in response to internal and external conditions - There is interest in using transcriptomes for predicting phenotypic traits and disorders - It is unknown how many samples are required for accurate phenotype prediction - We created 12 massive RNA-seq datasets and corresponding phenotypic data and used these to create prediction models - Performance curves from our models provide insight into the required number of samples for accurate prediction
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Harnessing genetic diversity in the USDA pea (Pisum sativum L.) germplasm collection through genomic prediction 93%
- A recently formed triploid Cardamine insueta inherits leaf vivipary and submergence tolerance traits of parents 92%
- Using local convolutional neural networks for genomic prediction 91%
Similar papers in this journal
- Multi-trait random regression models increase genomic prediction accuracy for a temporal physiological trait derived from high-throughput phenotyping 94%
- Comparative transcriptomic analysis reveals key components controlling spathe color in Anthurium andraeanum (Hort.) 92%
- Targeting Enhanced Digestibility: Prioritizing Low Pith Lignification to Complement low p-Coumaric Acid content as environmental stress intensity increase 92%
Similar papers in this journal
- Temporally resolved growth patterns reveal novel information about the polygenic nature of complex quantitative traits 94%
- Systematic analysis of 1,298 RNA-Seq samples and construction of a comprehensive soybean (Glycine max) expression atlas 94%
- An atlas of the Norway spruce needle seasonal transcriptome 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.