The utility of single-cell RNA sequencing data in predicting plant metabolic pathway genes
Ma, J.; Zou, L.; Wang, Z.; Wang, X.; Zuo, X.; Wang, F.; Wang, Z.; Li, Z.; Li, L.; Wang, P.
Show abstract
It is an ever challenging task to make genome-wide predictions for plant metabolic pathway genes (MPGs) encoding enzymes that catalyze the biosynthesis of plant natural products. Here, starting from 1,129 benchmark MPGs that have experimental evidence in Arabidopsis thaliana, we investigate the utilities of single-cell RNA sequencing (scRNA-seq) data--a recently arisen omics data that has been used in several other fields--in predicting MPGs using five machine learning (ML) algorithms that support multi-label tasks. Compared with traditional bulk RNA-seq data, scRNA-seq data lead to different but comparable co-expression networks among MPGs within metabolic classes, and significantly higher prediction accuracy of MPGs into classes. Prediction accuracy for individual metabolic classes is not associated with the co-expression network tightness, but correlated with the number of MPGs within each class, indicating that including more benchmark genes in the future will improve the MPG prediction. Splitting the RNA-seq data into genetic background/condition or tissue-specific subsets can improve the gene co-expression network tightness and MPG prediction accuracy for some classes; scRNA-seq-based models still outperform bulk RNA-seq-based models for most classes when corresponding subsets are used. In addition, deep learning approaches outperform classical machine learning approaches; approaches implemented in an ensembled workflow AutoGluon tend to have severe overfitting issues potentially due to the relative scarcity of benchmark MPGs within classes. Our results demonstrate the superiority of scRNA-seq data over bulk RNA-seq data in predicting MPGs into metabolic classes, and propose that scRNA-seq data should be included in the future to advance the identification of plant MPGs.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Chromosome-level genome assembly of Torreya grandis provides insights into the origin and evolution of gymnosperm-specific sciadonic acid biosynthesis 95%
- A Day in the Life of Arabidopsis: 24-Hour Time-lapse Single-nucleus Transcriptomics Reveal Cell-type specific Circadian Rhythms 95%
- Time-resolved oxidative signal convergence across the algae-embryophyte divide 94%
Similar papers in this journal
- Unique and distinct identities and functions of leaf phloem cells revealed by single cell transcriptomics 94%
- CAM evolution is associated with gene family expansion in an explosive bromeliad radiation 93%
- Transcriptional activation of auxin biosynthesis drives developmental reprogramming of differentiated cells 93%
Similar papers in this journal
- The Taxus genome provides insights into paclitaxel biosynthesis 95%
- An atlas of plant full-length RNA reveals tissue-specific and evolutionarily-conserved regulation of poly(A) tail length 94%
- Comparative transcriptomics in ferns reveals key innovations and divergent evolution of secondary cell wall 94%
Similar papers in this journal
- GeneWalk identifies relevant gene functions for a biological context using network representation learning 94%
- NetAct: a computational platform to construct core transcription factor regulatory networks using gene activity 93%
- Multi-omics analysis reveals the molecular response to heat stress in a "red tide" dinoflagellate 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.