Telomere-to-telomere assemblies reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by helitrons
Kainer, D.; Martin, S.; Hopp, D.; Mosher, S.; Tschaplinski, T. J.; Hyatt, P. D.; Martin, M. Z.; Leboldus, J. M.; Sondreli, K. L.; Busby, P. E.; Shu, M.; Barry, K.; Schmutz, J.; Furches, A.; Zhao, N.; Jacobson, D.; Chen, J.-G.; Pavicic, M.; Ranjan, P.; Muchero, W.; Tuskan, G. A.; Garvin, M. R.
Show abstract
BackgroundThe model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases (KCS), which has been difficult to confirm based on short-read assemblies. New long-read sequencing provides opportunities to develop telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed in traditional analyses. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. ResultsOur analysis of 78 telomere-to-telomere long-read haplotypes identified more than twice as many KCS genes as previously reported, along with numerous intragenic non-synonymous substitutions. Random forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a helitron transposon. A phylogenetic analysis indicates they are the evolutionary mechanism for generating KCS tandem arrays. ConclusionsLong-read sequencing and telomere-to-telomere assembles revealed large-effect loci critical to genetic studies that are unattainable from short-reads. These approaches also have the potential to reveal novel insights into genome structure and function, such as the helitrons identified here. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genome assembly and characterization of a complex zfBED-NLR gene-containing disease resistance locus in Carolina Gold Select rice with Nanopore sequencing 95%
- Genetic and environmental influences on the distributions of three chromosomal drive haplotypes in maize 94%
- The VIL gene CRAWLING ELEPHANT controls maturation and differentiation in tomato via polycomb silencing 94%
Similar papers in this journal
- MYB5a/NEGAN activates petal anthocyanin pigmentation and shapes the MBW regulatory network in Mimulus luteusvar. variegatus 94%
- Ozone sensitivity of diverse maize genotypes is associated with differences in gene regulation, not gene content 94%
- Natural variation modifies centromere proximal meiotic crossover frequency and segregation distortion in Arabidopsis thaliana 94%
Similar papers in this journal
- Homogeneity among glyphosate-resistant Amaranthus palmeri in geographically distant locations 95%
- Contrasting a reference cranberry genome to a crop wild relative provides insights into adaptation, domestication, and breeding 95%
- Distinctiveness of genes contributing to growth of Pseudomonas syringae in diverse host plant species 94%
Similar papers in this journal
- Genomic characterization of a nematode tolerance locus in sugar beet 95%
- Large scale genomic rearrangements in selected Arabidopsis thaliana T-DNA lines are caused by T-DNA insertion mutagenesis 94%
- Differential gene expression associated with a floral scent polymorphism in the evening primrose Oenothera harringtonii (Onagraceae) 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.