Back

Predicting plant traits using large conglomerate RNA-seq datasets

Hadish, J. A. A.; Honaas, L.; Ficklin, S.

2024-12-13 systems biology
10.1101/2024.12.09.627626 bioRxiv
Show abstract

RNA-seq datasets offer potential for predicting phenotypic traits, but the optimal dataset size for reliable predictions remains unclear. To explore this question, we compiled large-scale transcriptomic datasets across 12 plant species and used them to predict the phenotypic variables of tissue type and age. Predictions were made using random forest models. We created these models using an increasing number of samples to create performance curves. These curves show that only a few hundred samples are required to achieve maximum accuracy for tissue classification, while predicting age demands a few thousand samples. Acceptable prediction accuracy can be achieved at even lower numbers of samples. Our findings provide a benchmark for designing transcriptomic studies aimed at phenotype prediction and highlight the differing complexities involved in predicting simple versus more complex traits. CORE IDEAS- Plant transcriptomes are highly dynamic in response to internal and external conditions - There is interest in using transcriptomes for predicting phenotypic traits and disorders - It is unknown how many samples are required for accurate phenotype prediction - We created 12 massive RNA-seq datasets and corresponding phenotypic data and used these to create prediction models - Performance curves from our models provide insight into the required number of samples for accurate prediction

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.