DNA Research
◐ Oxford University Press (OUP)
All preprints, ranked by how well they match DNA Research's content profile, based on 26 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Shirasawa, K.; Arimoto, R.; Hirakawa, H.; Ishimorai, M.; Ghelfi, A.; Miyasaka, M.; Endo, M.; Kawabata, S.; Isobe, S.
Show abstract
Eustoma grandiflorum (Raf.) Shinn., is an annual herbaceous plant native to the southern United States, Mexico, and the Greater Antilles. It has a large flower with a variety of colors and an important flower crop. In this study, we established a chromosome-scale de novo assembly of E. grandiflorum by integrating four genomic and genetic approaches: (1) Pacific Biosciences (PacBio) Sequel deep sequencing, (2) error correction of the assembly by Illumina short reads, (3) scaffolding by chromatin conformation capture sequencing (Hi-C), and (4) genetic linkage maps derived from an F2 mapping population. The 36 pseudomolecules and unplaced 64 scaffolds were created with total length of 1,324.8 Mb. Full-length transcript sequencing was obtained by PacBio Iso-Seq sequencing for gene prediction on the assembled genome, Egra_v1. A total of 36,619 genes were predicted on the genome as high confidence HC) genes. Of the 36,619, 25,936 were annotated functions by ZenAnnotation. Genetic diversity analysis was also performed for nine commercial E. grandiflorum varieties bred in Japan, and 254,205 variants were identified. This is the first report of the construction of reference genome sequences in E. grandiflorum as well as in the family Gentianaceae.
Yang, J.; Zhang, R.; Ma, Y.; Ma, Y.; Sun, W.
Show abstract
The tree species Firmiana major was once dominant in the savanna vegetation of the arid hot valleys of southwest China, but was considered extinct in the wild in 1998. After eight small populations were relocated by thorough investigations between 2018 and 2020, the species was subsequently recognized as a Plant Species of Extremely Small Populations (PSESP) in China in need of urgent rescue. Moreover, due to severe human disturbance, other species in the tropical woody genus Firmiana are also endangered, and the species in this genus have almost all been listed as second-class National Protected Wild Plants in China. In order to guide future research into the conservation of this group, we present here the high-quality genome assembly of F. major. This is the first genome assembly in the genus Firmiana, and is 1.4 Gb in size. The assembly consists of 1.18 Gb repetitive sequences, 37,673 annotated genes and 31,965 coding genes.
Liu, S.; Wang, Z.; Shi, T.; Dan, X.; Zhang, Y.; Liu, J.; Wang, J.
Show abstract
High-quality reference genomes for several species have promoted breeding and functional studies of poplar trees. By resequencing numerous accessions of these and closely related species, single nucleotide polymorphisms (SNPs) and small insertion/deletions (InDels) have been identified to assist in clarifying local adaptation and phenotypic diversification. A chromosome-level genome assembly for P. adenopoda was assembled based on Illumina and PacBio sequencing platforms, facilitated by Hi-C technology. The assembled genome size was about 383 Mb, with 99.70% of the contigs anchored to 19 pseudo-chromosomes, and a total of 33,505 protein-coding genes were annotated. This high-quality genome provided the genomic basis for the subsequent detection of various variants.
Isobe, S.; Fujii, H.; Shirasawa, K.; Kawahara, Y.; Endo, T.; Shimada, T.
Show abstract
Citrus, a member of the Rutaceae family, is a widely cultivated crop with numerous cultivars. In Japan, citrus fruits account for a significant portion of agricultural production. Although several new citrus varieties have been developed through conventional breeding programs, satsuma mandarin remains the dominant cultivar. In this study, chromosome-scale and haploid-resolved reference genome sequences of satsuma mandarin (Citrus unshiu Marc) and its parental varaieties, kishu mandarin (C. kinokuni hort. ex Tanaka) and kunenbo mandarin (C. nobilis Lour. var. kunip Tanaka) were generated using long-read sequencing and Hi-C technologies. The comparison of haploid and unphased genomes revealed structural differences between them, indicating distinct regions in each haploid. In addition, genetic linkage maps were constructed, and genetic and physical distances were compared. The results showed variations in polymorphism density across different regions of the chromosomes. Together, the obtained results provide valuable insights into the genomic characteristics and structural variations of satsuma mandarin and related citrus varieties. These insights will lead to the further elucidation and improvement of citrus cultivars through genome breeding strategies.
Shirasawa, K.; Akita, Y.; Mizunoe, Y.; Takamura, T.
Show abstract
Cyclamen is an economically important ornamental plant widely cultivated for its diverse floral characteristics and adaptation to cool climates. Despite its horticultural significance, genomic resources for this species remain limited, hindering molecular studies and genomics-assisted breeding. Here, we report the first highly contiguous nuclear genome assembly of C. persicum generated using high-fidelity long-read sequencing. The assembled genome spans 1.48 Gb, consisting of 126 contigs with an N50 length of 52.3 Mb. Telomeric repeat analysis identified eight contigs containing telomeric sequences at both ends, suggesting the presence of near-complete chromosome assemblies. Genome completeness assessment using BUSCO indicated 98.1% completeness. Repetitive sequences occupied 82.9% of the assembly, with long terminal repeat retrotransposons accounting for 42.1% of the genome. A total of 40,223 protein-coding genes were predicted, with a complete BUSCO score of 95.7%. Comparative orthogroup analysis with five representative eudicot species identified 430 orthogroups specific to C. persicum and 363 orthogroups shared exclusively between C. persicum and Primula kwangtungensis, indicating the presence of both lineage-specific and Primulaceae-conserved gene families. These findings provide critical insights into gene family evolution within Primulaceae and establish an essential comparative framework for future genomic studies. The genome resource presented here provides an invaluable foundation for investigating genome evolution, gene function, and trait-associated loci in cyclamen, effectively facilitating molecular breeding and genetic improvement in this ornamental species.
Chen, L.; Jiang, Y.; Shi, T.; Jia, C.; Long, Z.; Zhang, X.; Sang, Y.; Liu, J.; Wang, J.
Show abstract
Despite the population structure and genetic differentiation of P. davidiana has been reported, little is known about the P. davidiana genome characterization with the genomes assembled at the chromosome level. As one of the most widely distributed and ecologically important tree species in China, P. davidiana is an excellent resource to understand population adaption to climate and environment change. Here we present a high-quality assembly based on long-read sequences and Hi-C data. This assembly is assembled into 19 contiguous chromosomes which provides a powerful tool for future association studies.
Fujiwara, K.; Toyoda, A.; Biswa, B. B.; Kishida, T.; Tsuruta, M.; Nakamura, Y.; Kimura, N.; Kawamoto, S.; Sato, Y.; Katsuki, T.; Sakura 100 Genome Consortium, ; Koide, T.
Show abstract
The Oshima cherry (Cerasus speciosa), which is endemic to Japan, has significant cultural and horticultural value. In this study, we present a near complete telomere-to-telomere genome assembly for C. speciosa, derived from the old growth "Sakurakkabu" tree on Izu Oshima Island. Using Illumina short-read, PacBio long-read, and Hi-C sequencing, we constructed a 269.3 Mbp genome assembly with a contig N50 of 32.0 Mbp. We examined the distribution of repetitive sequences in the assembled genome and identified regions that appeared to be centromeric. Detailed structural analysis of these putative centromeric regions revealed that the centromeric regions of C. speciosa comprised repetitive sequences with monomer lengths of 166 or 167 bp. Comparative genomic analysis with Prunus sensu lato genome revealed structural variations and conserved syntenic regions. This high-quality reference genome provides a crucial tool for studying the genetic diversity and evolutionary history of Cerasus species, facilitating advancements in horticultural research and the preservation of this iconic species.
Long, Z.; Sang, Y.; Feng, J.; Shi, T.; Dan, X.; Zhang, Y.; Liu, J.; Wang, J.
Show abstract
Despite widespread biodiversity loss, our understanding of how species and populations will respond to accelerated climate change remains limited. In this study, we predict the evolutionary responses of Populus lasiocarpa, a key alpine forest tree species primarily found in the mountainous regions of a global biodiversity hotspot, to climate change. We accomplish this by generating and integrating a new reference genome for P. lasiocarpa, re-sequencing data for 200 samples, and gene expression profiles for leaf and root tissue following exposure to heat and waterlogging. Analyses of the re-sequencing data indicate that demographic dynamics, divergent selection, and long-term balancing selection have shaped and maintained genetic variation within and between populations over historical timescales. In examining genomic signatures of contemporary climate adaptation, we found that haplotype blocks, characterized by inversion polymorphisms that suppress recombination, play a crucial role in clustering environmentally adaptive variations. Comparison of evolved and plastic gene expression show that genes with expression plasticity generally align with evolved responses, highlighting the adaptive role of plasticity. Lastly, we incorporated local adaptation, migration, genetic load, and plasticity responses into our predictions of population-level climate change risks. Our findings reveal that western populations, primarily distributed in the Hengduan Mountains--a region known for its environmental heterogeneity and significant biodiversity--are the most vulnerable to climate change and should be prioritized for conservation and management. Overall, our study advances understanding of the relative roles of long-term natural selection, local environmental adaptation, and immediate plasticity responses in driving evolutionary adaptation to climate change in keystone species.
Li, W.; Zhu, X.-G.; Zhang, Q.-J.; Li, K.; Zhang, D.; Shi, C.; Gao, L.-Z.
Show abstract
Mango (Mangifera indica), a member of the family Anacardiaceae, is one of the worlds most popular tropical fruits. Here we sequenced the variety, "Hong Xiang Ya", and generated a 371.6-Mb mango genome assembly with 34,529 predicted protein-coding genes. Aided with the published genetic map, for the first time, we assembled the M. indica genome to the chromosomes, and finally about 98.77% of the genome assembly was anchored to 20 pseudo-chromosomes. The availability of the chromosome-length genome assembly of M. indica will provide novel insights into genome evolution, understand the genetic basis of specialized phytochemical composites relevant to fruit quality, and enhance allele mining in genomics-assisted breeding for mango genetic improvement.
zhang, y.; Wang, D.; Zhao, R.; Li, S.; Zheng, X.; Hu, G.
Show abstract
Rana dybowskii is distributed across Northeast Asian and represents a valuable medical resource. A high-quality assembly of the genome has not yet been reproted. This species has 2n=24 chromosomes, but a huge genome size that estimated at 3.5 ~4.6 Gb in the previous studies. The relatively large chromosome size, exceeding hundreds of megabases, may result in difficulties of obtaining a complete chromosome level genome. Here, we constructed a chromosome-level genome assembly of R. dybowskii by integrating PacBio HiFi long-read sequencing for de novo assembly and CiFi (3C coupled with HiFi sequencing) for scaffolding. The final assembly consists of 12 chromosomes with a total of 3.95 Gb and a scaffold N50 length of 455 Mb. BUSCO assessment using the tetrapoda_odb12 database identified 94.2% complete and 0.5% fragmented orthologs, suggesting a high level of completeness of the assembly. Genomic annotation revealed that repetitive sequences comprise over 53% of the assembly, with retroelements and DNA transposons accounting for 22% and 25%, respectively. A total of 43,999 protein-coding genes were predicted with the assistance of RNA-seq reads from four tissues (muscle, eye, testis and skin). This high-quality chromosome-level reference genome provides a valuable genomic resource for advancing genetic studies of the species.
Yoon, U.-H.; Cao, Q.; Shirasawa, K.; Zhai, H.; Lee, T.-H.; Tanaka, M.; Hirakawa, H.; Hahn, J.-H.; Wang, X.; Kim, H. S.; Tabuchi, H.; Zhang, A.; Kim, T.-H.; Nagasaki, H.; Xiao, S.; Okada, Y.; Jeong, J. C.; Nagano, S.; Shin, Y.; Lee, H.-U.; Park, S.-U.; Lee, S. J.; Lee, K.; Yang, J.-W.; Ahn, B. O.; Ma, D.; Takahata, Y.; Kwak, S.-S.; Liu, Q.; Isobe, S.
Show abstract
Sweetpotato (Ipomoea batatas (L.) Lam) is the worlds seventh most important food crop by production quantity. Cultivated sweetpotato is a hexaploid (2n = 6x = 90), and its genome (B1B1B2B2B2B2) is quite complex due to polyploidy, self-incompatibility, and high heterozygosity. Here we established a haploid-resolved and chromosome-scale de novo assembly of autohexaploid sweetpotato genome sequences. Before constructing the genome, we created chromosome-scale genome sequences in I. trifida using a highly homozygous accession, Mx23Hm, with PacBio RSII and Hi-C reads. Haploid-resolved genome assembly was performed for a sweetpotato cultivar, Xushu18 by hybrid assembly with Illumina paired-end (PE) and mate-pair (MP) reads, 10X genomics reads, and PacBio RSII reads. Then, 90 chromosome-scale pseudomolecules were generated by aligning the scaffolds onto a sweetpotato linkage map. De novo assemblies were also performed for chloroplast and mitochondrial genomes in I. trifida and sweetpotato. In total, 34,386 and 175,633 genes were identified on the assembled nucleic genomes of I. trifida and sweetpotato, respectively. Functional gene annotation and RNA-Seq analysis revealed locations of starch, anthocyanin, and carotenoid pathway genes on the sweetpotato genome. This is the first report of chromosome-scale de novo assembly of the sweetpotato genome. The results are expected to contribute to genomic and genetic analyses of sweetpotato.
Hu, Y.; Huang, Y.; Yong, Y.; Shang, E.; Zhang, B.; Sui, Z.
Show abstract
As an important cultivated red alga, Gracilariopsis lemaneiformis has great economic and ecological value. However, its existing genome assembly is highly fragmented and inadequately annotated. In this study, we constructed the first high-quality chromosome-level genome of Gp. lemaneiformis using PacBio long reads, Illumina short reads and Hi-C sequencing data. The assembled genome was approximately 86.66 Mb and the assembled sequences were anchored to 28 pseudo-chromosomes with lengths ranging from 1.70 to 7.81 Mb. 99.91% of the PacBio reads could be mapped to our assembly. In total, 8,664 genes were annotated, and the repeat elements identified in Gp. lemaneiformis constituted 65.04% of the whole genome, including 2.24% tandem repeat sequences and 62.81% interspersed repeats. We also established a high-evidence phylogenetic tree from 19 representative algae species, with the main aim to calculate their divergence times. This high-quality genome of Gp. lemaneiformis provides a crucial foundation for understanding genetic characteristics, investigating the genomic evolution, and facilitating molecular breeding.
Zhang, F.; Yang, Y.-h.; Li, W.; Shi, C.; Zhu, X.-g.; Gao, L.-z.
Show abstract
Oryza granulata Nees et Arn. ex Watt, a diploid wild rice (GG genome), possesses exceptional shade tolerance and is a key genetic resource for rice improvement. However, previous genome assemblies lacked continuity and completeness. Here we present a chromosome-scale reference genome of O. granulata using PacBio SMRT (113x), Hi-C (95x), and Illumina sequencing. The final assembly is ~764.24 Mb, with a scaffold N50 of ~59.32 Mb, and ~96.47% of the sequence anchored to 12 chromosomes. BUSCO completeness is ~98.6%. We annotated ~42,064 protein-coding genes, of which ~95.39% were functionally annotated, along with ~73.46% repetitive elements. The genome assembly and raw sequencing data are available at NGDC (PRJCA061980), NGDC GSA (CRA068332), and NGDC GWH (GWHISVE00000000.1). This high-quality genome will serve as a fundamental resource for evolutionary genomics, conservation biology, and breeding of shade-tolerant rice cultivars.
Hosaka, A.
Show abstract
Transposable Elements (TEs) are major components of the genome. To understand their function and evolution, it is necessary to identify active TEs from a diverse range of organisms. Here, I report the genome of the Nishikigoi, an ornamental fish derived from the Common carp, and the novel approach to detecting active TE candidates. I constructed a chromosome-scale assembly using long-read sequencing and Hi-C methods. It revealed that Nishikigoi has Robertsonian-like chromosomal translocations not seen in Common carp. I also found that Nishikigoi has a significantly different genetic background from Common carp, reflecting the intensive breeding history. Furthermore, by focusing on Runs of Homozygosity (ROH) islands in the Nishikigoi genome and analyzing structural variations with long-read sequencing, I identified several active TE candidates. This study not only revealed the unique genetic features of Nishikigoi but also demonstrated the potential for a novel approach in the search for active TEs.
Jia, S.; Wang, G.; Liu, G.; Qu, J.; Zhao, B.; Jin, X.; Zhang, L.; Yin, J.; Liu, C.; Shan, G.; Wu, S.; Song, L.; Liu, T.; Wang, X.; Yu, J.
Show abstract
The red algae Kappaphycus alvarezii is the most important aquaculture species in Kappaphycus, widely distributed in tropical waters, and it has become the main crop of carrageenan production at present. The mechanisms of adaptation for high temperature, high salinity environments and carbohydrate metabolism may provide an important inspiration for marine algae study. Scientific background knowledge such as genomic data will be also essential to improve disease resistance and production traits of K. alvarezii. 43.28 Gb short paired-end reads and 18.52 Gb single-molecule long reads of K. alvarezii were generated by Illumina HiSeq platform and Pacbio RSII platform respectively. The de novo genome assembly was performed using Falcon_unzip and Canu software, and then improved with Pilon. The final assembled genome (336 Mb) consists of 888 scaffolds with a contig N50 of 849 Kb. Further annotation analyses predicted 21,422 protein-coding genes, with 61.28% functionally annotated. Here we report the draft genome and annotations of K. alvarezii, which are valuable resources for future genomic and genetic studies in Kappaphycus and other algae.
Fan, Y.; Sahu, S. K.; Yang, T.; mu, w.; wei, j.; Cheng, L.; Yang, J.-l.; Xu, X.; Liu, X.; Mu, R.-c.; Liu, J.; Zhao, J.-m.; Zhao, Y.-x.; Liu, H.
Show abstract
The Averrhoa carambola is commonly known as star fruit because of its peculiar shape and its fruit is a rich source of minerals and vitamins. It is also used in traditional medicines in countries like India, China, the Philippines, and Brazil for treating various ailments such as fever, diarrhea, vomiting, and skin disease. Here we present the first draft genome of the Oxalidaceae family with an assembled genome size of 470.51 Mb. In total, 24,726 protein-coding genes were identified and 16,490 genes were annotated using various well-known databases. The phylogenomic analysis confirmed the evolutionary position of the Oxalidaceae family. Based on the gene functional annotations, we also discovered the enzymes possibly involved in the important nutritional pathways in star fruit genome. Overall, being the first sequenced genome in the Oxalidaceae family, the data provides an essential resource for the nutritional, medicinal, and cultivational studies for this economically important star-fruit plant.
Morikami, K.; Tanizawa, Y.; Yagura, M.; Sakamoto, M.; Kawamoto, S.; Nakamura, Y.; Yamaguchi, K.; Shigenobu, S.; Naruse, K.; Ansai, S.; Kuraku, S.
Show abstract
Medaka, a group of small, mostly freshwater fishes in the teleost order Beloniformes, includes the rice fish Oryzias latipes, which is a prominent model organism for diverse biological fields. Chromosome-scale genome sequences of the Hd-rR strain of this species were obtained in 2007, and its improved version has facilitated various genome-wide studies. However, despite its widespread utility, omics data for O. latipes are dispersed across various public databases and lack a centralized platform. To address this, the medaka section of the National Bioresource Project (NBRP) of Japan established a genome informatics team in 2022 tasked with providing versatile in silico solutions for bench biologists. This initiative led to the launch of MedakaBase (https://medakabase.nbrp.jp), a web server that enables gene-oriented analysis including exhaustive sequence similarity searches. MedakaBase also provides genome-wide browsing of diverse datasets, including tissue-specific transcriptomes and intraspecific genomic variations, integrated with gene models from different sources. Additionally, the platform offers gene models optimized for single-cell transcriptome analysis, which often requires coverage of the 3' untranslated region (UTR) of transcripts. Currently, MedakaBase provides genome-wide data for seven Oryzias species, including original data for O. mekongensis and O. luzonensis produced by the NBRP team. This article outlines technical details behind the data provided by MedakaBase.
Masuoka, Y.; Jouraku, A.; Kuwazaki, S.; Yoshiyama, M.; Horigane-Ogihara, M.; Maeda, T.; Suzuki, Y.; Bono, H.; Kimura, K.; Yokoi, K.
Show abstract
Honey bees are important for agriculture (e.g., pollination and honey production). Additionally, honey bees are an important insect model species, especially as model social insects. The Japanese honey bee, Apis cerana japonica (a subspecies of the Asian honey bee, Apis cerana), is a Japanese domestic honey bee, which has several subspecies-specific traits. We previously constructed the draft genome sequence data of A. cerana japonica, but it needed to be improved considering the use of the genome sequence data for genome structural analysis and repetitive region analysis, as well as the availability of chromosome-level genome data of A. mellifera and A. cerana. In this study, we constructed the improved A. cerana japonica genome data and new gene set data with functional annotations. The constructed genome data, including 16 pseudochromosomes, was found to be highly contiguous and complete, and the gene set data covered most of the core genes in the BUSCO database. Thus, the constructed genome and gene set data have become more suitable as the reference data of A. cerana japonica.
Gaikwad, K.; Ramakrishna, G.; Srivastava, H.; Saxena, S.; Kaila, T.; Tyagi, A.; Sharma, P.; Sharma, S.; Sharma, R.; Mahla, H.; SV, A. M.; Solanke, A.; Kalia, P.; Rao, A.; Rai, A.; Sharma, T.; Singh, N.
Show abstract
Clusterbean (Cyamopsis tetragonoloba (L.) Taub.), also known as Guar is a widely cultivated dryland legume of Western India and parts of Africa. Apart from being a vegetable crop, it is also an abundant source of a natural hetero-polysaccharide called guar gum or galactomannan which is widely used in cosmetics, pharmaceuticals, food processing, shale gas drilling etc. Here, for the first time we are reporting a chromosome-scale reference genome assembly of clusterbean, from a high galactomannan containing popular guar cultivar, RGC-936, by combining sequenced reads from Illumina, 10x Chromium and Oxford Nanopore technologies. The initial assembly of 1580 scaffolds with an N50 value of 7.12 Mbp was generated. Then, the final genome assembly was obtained by anchoring these scaffolds to a high density SNP map. Finally, a genome assembly of 550.31 Mbp was obtained in 7 pseudomolecules corresponding to 7 chromosomes with a very high N50 of 78.27 Mbp. We finally predicted 34,680 protein-coding genes in the guar genome. The high-quality chromosome-scale cluster bean genome assembly will facilitate understanding of the molecular basis of galactomannan biosynthesis and aid in genomics-assisted breeding of superior cultivars.
Li, H.; Rehman, S. u.; Song, R.; Qiao, L.; Hao, X.; Zhang, J.; Li, K.; Hou, L.; Hu, W.; Wang, L.; Chen, S.
Show abstract
Wild relatives of wheat are valuable sources for enhancing the genetic diversity of common wheat. Aegilops comosa, an annual diploid species with an MM genome constitution, possesses numerous agronomically valuable traits that can be exploited for wheat improvement. In this study, we report a chromosome-level genome assembly of Ae. comosa accession PI 551049, generated using PacBio high-fidelity (HiFi) reads and high-throughput chromosome conformation capture (Hi-C) data. The assembly spans 4.47 Gb, featuring a contig N50 of 23.59 Mb and a scaffold N50 of 619.05 Mb. A total of 39,063 gene models were annotated through a combination of homoeologous proteins, Iso-Seq, and RNA-Seq data. Comparative genome analysis revealed a terminal intrachromosomal translocation in chromosome 2M of Ae. comosa (and Ae. umbellulata) compared to its homoeologous chromosomes in other diploid wheat species. Phylogenetic analysis showed a close relationship between Ae. comosa and Ae. umbellulata. This newly constructed reference genome of Ae. comosa will serve as an important genomic resource for comparative genomic studies and the cloning of agriculturally important genes.