Back

SoyFGB v2.0: a unique access to variations of Chinese Soybean Gene Bank (CNSGB) germplasm

Zheng, T.; Li, Y.; Li, Y.; Zhang, S.; Wang, C.; Zhang, F.; Zhang, L.; Wu, X.; Tian, Y.; Jiang, S.; Xu, J.; Qiu, L.

2021-12-30 genomics
10.1101/2021.12.28.474253 bioRxiv
Show abstract

In Chinese National Soybean GeneBank (CNSGB), we have collected more than 30,000 soybean accessions. However, data sharing for soybean remains an especially sensitive question, and how to share the genome variations within rule frame has been bothering the soybean germplasm workers for a long time. Here we release a big data source named Soybean Functional Genomics & Breeding database (SoyFGB v2.0) (https://sfgb.rmbreeding.cn/), which embed a core collection of 2,214 soybean resequencing genome (2K-SG) from the CNSGB germplasm. This source presents a unique example which may help elucidating the following three major functions for multiple genome data mining with general interests for plant researchers. 1) On-line analysis tools are provided by the Analysis module for haplotype mining in high-throughput genotyped germplasms with different methods. 2) Variations for 2K-SG are provided in SoyFGB v2.0 by Browse module which contains two functions of SNP and InDel. Together with the Gene (SNP & InDel) function embedded in Search module, the genotypic information of 2K-SG for targeting gene / region is accessible. 3) Scaled phenotype data of 42 traits, including 9 quality and 33 quantitative traits are provided by SoyFGB v2.0. With the scaled-phenotype data search and seed request tools under a control list, the germplasm information could be shared without direct downloading the unpublished phenotypic data or information of sensitive germplasms. In a word, the mode of data mining and sharing underlies SoyFGB v2.0 may inspire more ideas for works on genome resources of not only soybean but also the other plants.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.