Computational Identification of Candidate Gene Families for Volatile Sulfur Compound Biosynthesis in Cannabis sativa Using Profile Hidden Markov Models
Maminakis, E.; Geffen, L.; Barbosa-Xavier, K.; Sharif, S.
Show abstract
Cannabis is well known for its pungent, skunk-like aroma. Recent chemical studies have identified prenylated and C6 volatile sulfur compounds as contributors to its skunky and citrus-like aromas, but the pathways that produce these compounds remain unknown. This gap limits efforts to explain variation in sulfur-aroma traits and to selectively enhance or reduce those traits. To address this gap, we used the known chemistry of sulfur-containing volatiles in Cannabis and characterized sulfur and volatile biosynthetic pathways in other plant species to select candidate enzyme groups. Because the GMO cultivar is anecdotally associated with a pronounced sulfurous aroma, reference protein sequences and profile hidden Markov models were used to search its version 1 (v1) primary high-confidence protein set of 55,790 sequences. These searches recovered 975 unique proteins. Sequence screening retained 941 candidates across 20 reporting categories; 939 contained all expected domains, while the two candidates assigned to the methionine gamma-lyase (MGL)-nearest category had no category-specific expected-domain rule. The largest reporting category comprised 359 proteins containing a cytochrome P450 domain, recovered through a search motivated by cytochrome P450 family 74 (CYP74) enzymes involved in oxylipin and plant volatile formation. Thirteen of these proteins were also recovered by at least one full-length CYP74 reference search. Other large reporting categories included 218 sugar-transferase, 83 glutathione-transferase, and 61 alcohol dehydrogenase candidates. Comparison with the Cannabis Expression Atlas linked 168 candidates to 128 annotated genes through 100%-identity amino-acid matches spanning at least 80% of each GMO v1 candidate protein. Twenty-nine genes were tissue-specific, including 13 root-specific and 6 trichome-specific genes. These results define candidates for biochemical testing and direct searches for additional enzymes acting upstream and downstream in Cannabis sulfur-volatile pathways.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Delineating the genetic regulation of cannabinoid biosynthesis during female flower development in Cannabis sativa 93%
- Hydroxyl carlactone derivatives are predominant strigolactones in Arabidopsis 92%
- High Spatial-Resolution Imaging Of The Dynamics Of Cuticular Lipid Deposition During Arabidopsis Flower Development 92%
Similar papers in this journal
- Catechol acetylglucose: A newly identified benzoxazinoid-regulated defensive metabolite in maize 92%
- Integrative metabolomics reveal the organisation of alkaloid biosynthesis in Daphniphyllum macropodum 92%
- Two bi-functional cytochrome P450 CYP72 enzymes from olive (Olea europaea) catalyze the oxidative C-C bond cleavage in the biosynthesis of secoxy-iridoids - flavor and quality determinants in olive oil 91%
Similar papers in this journal
- Conserved amino acid residues and gene expression patterns associated with the substrate preferences of the competing enzymes FLS and DFR 90%
- Phytohormone production by the arbuscular mycorrhizal fungus Rhizophagus irregularis 90%
- Microtranscriptome of contrasting sugarcane cultivars in response to aluminum stress 89%
Similar papers in this journal
- Unraveling Cross-Cellular Communication in Cannabis sativa Sex Determination: A Network Ontology Transcript Annotation (Nota) Analysis. 92%
- Influences of chemotype and parental genotype on metabolic fingerprints of tansy plants uncovered by predictive metabolomics. 92%
- Genetic Insights into Agronomic and Morphological Traits of Drug-Type Cannabis Revealed by Genome-Wide Association Studies 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.