Predicting Phenotypic Traits Using a Massive RNA-seq Dataset
Hadish, J. A.; Honaas, L. A.; Ficklin, S. P.
Show abstract
Transcriptomic data can be used to predict environmentally impacted phenotypic traits. This type of prediction is particularly useful for monitoring difficult-to-measure phenotypic traits and has become increasingly popular for monitoring high-value agricultural crops and in precision medicine. Despite this increase in popularity, little research has been done on how many samples are required for these models to be accurate, and which normalization should be used. Here we create a massive RNA-seq dataset from publicly available Arabidopsis thaliana data with corresponding measurements for age and tissue type. We use this dataset to determine how many samples are required for accurate model prediction and which normalization method is required. We find that Median Ratios Normalization significantly increases performance when predicting age. We also find that in the case of our dataset, only a few hundred samples are required to predict tissue types, and only a few thousand samples are necessary to accurately predict age. Researchers should consider these results when choosing the number of samples in a transcriptomic experiment and during data-processing. Author SummaryLarge datasets have become ubiquitous in both research and industry, with thousands and sometimes millions of samples being collected for a single project. In biology a prominent new technology is RNA-seq, which can be used to measure the expression level of thousands of genes for a single sample. These measurements are used for a variety of downstream applications, including predicting phenotypic traits (i.e. height, disease, etc.). A number of experiments have attempted to use RNA-seq data to make phenotype predictions with varying success. This is partially due to the small sample size of their experiments. RNA-seq datasets are currently relatively small--only a dozen to a few hundred samples--due to the cost per sample. This is expected to change as the cost of sequencing decreases. In this paper we create a massive conglomerate RNA-seq dataset from publicly available Arabidopsis thaliana RNA-seq data. We use this dataset to determine how many samples are required to accurately predict plant age and tissue type using machine learning models. We also explore the best way to normalize large datasets. Our results show the potential of massive RNA-seq datasets, and can be used to inform experimental design for phenotype prediction.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sparse Multitask group Lasso for Genome-Wide Association Studies 93%
- Modeling cross-regulatory influences on monolignol transcripts and proteins under single and combinatorial gene knockdowns in Populus trichocarpa 92%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 92%
Similar papers in this journal
- RWRtoolkit: multi-omic network analysis using random walks on multiplex networks in any species 94%
- scShapes: A statistical framework for identifying distribution shapes in single-cell RNA-sequencing data 93%
- Multi-Dimensional Machine Learning Approaches for Fruit Shape Recognition and Phenotyping in Strawberry 93%
Similar papers in this journal
- FINDER: An automated software package to annotate eukaryotic genes from RNA-Seq data and associated protein sequences 93%
- On the importance of data transformation for data integration in single-cell RNA sequencing analysis 93%
- Probabilistic modeling methods for cell-free DNA methylation based cancer classification 92%
Similar papers in this journal
- pISA-tree - a data management framework for life science research projects using a standardised directory tree 90%
- Cultivar-specific transcriptome and pan-transcriptome reconstruction of tetraploid potato 90%
- CF-Seq, An Accessible Web Application for Rapid Re-Analysis of Cystic Fibrosis Pathogen RNA Sequencing Studies 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.