A versatile resource of 1500 diverse wild and cultivated soybean genomes for post-genomics research
Zhang, H.; Jiang, H.; Hu, Z.; Song, Q.; An, Y.-q. C.
Show abstract
With the advance of next-generation sequencing technologies, over 15 terabytes of raw soybean genome sequencing data were generated and made available in the public. To develop a consolidated, diverse, and user-friendly genomic resource to facilitate post-genomic research, we sequenced 91 highly diverse wild soybean genomes representing the entire US collection of wild soybean accessions to increase the genetic diversity of the sequenced genomes. Having integrated and analyzed the sequencing data with the public data, we identified and annotated 32 million single nucleotide polymorphisms (32mSNPs) with a resolution of 30 SNPs/kb and 12 non-synonymous SNPs/gene in 1,556 accessions (1.5K). Population structure analysis showed that the 1.5K accessions represent the genetic diversity of the 20,087 (20K) soybean accessions in the U.S. collection. Inclusion of wild soybean genomes significantly increased the genetic diversity and shorten linkage disequilibrium distance in the panel of soybean accessions. We identified a collection of paired accessions sharing the highest genomic identity between the 1.5K and 20K accessions as genomically "equivalent" accessions to maximize the use of the genome sequences. We demonstrated that the 32mSNPs in the 1.5K accessions can be effectively used for in-silico genotyping, discovering trait QTL, gene alleles/mutations, identifying germplasms containing beneficial allele and domestication selection of trait alleles. We made the 32mSNPs and 1.5K accessions with detailed annotation available at SoyBase and Ag Data Commons. The dataset could serve as a versatile resource to release the potential of the huge amount of genome sequencing data for a variety of postgenomic research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Degenerate oligonucleotide primer MIG-seq: an effective PCR-based method for high-throughput genotyping 97%
- ALI-1, candidate gene of B1 locus, is associated with awn length and grain weight in common wheat 96%
- Assembly, comparative analysis, and utilization of a single haplotype reference genome for soybean 96%
Similar papers in this journal
- Haplotype mapping uncovers unexplored variation in wild and domesticated soybean at the major protein locus cqProt-003 95%
- Genetic characterization of cucumber genetic resources in the NARO Genebank indicates their multiple dispersal trajectories to the East 95%
- Identification of a Major Locus for Flowering Pattern Sheds Light on Plant Architecture Diversification in Cultivated Peanut 95%
Similar papers in this journal
- CRISPR-based editing of the ω- and γ-gliadin gene clusters reduces wheat immunoreactivity without affecting grain protein quality 95%
- Heritable temporal gene expression patterns correlate with metabolomic seed content in developing hexaploid oat seed 94%
- A Citrullus genus super-pangenome reveals extensive variations in wild and cultivated watermelons and sheds light on watermelon evolution and domestication 94%
Similar papers in this journal
- Identification of QTLs for Internode Length and Diameter Associated with Lodging Resistance in Rice 96%
- QTL mapping for seed morphology using the instance segmentation neural network in Lactuca spp. 95%
- Chromosome-level genome assemblies for two quinoa inbred lines from northern and southern highlands of Altiplano where quinoa originated 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.