VN1K: a genome graph-based and function-driven multi-omics and phenomics resource for the Vietnamese population
Tran, T. T. H.; Hoang, T. H.; Tran, M. H.; Nguyen, N. T.; Nguyen, D. T.; Pham, T.; Pham, T. M.; Nguyen, N. N.; Vu, G. M.; Duong, V. C.; Vu, Q. T.; Nguyen, T. K.; Nguyen, T. M.; Vu, H. Q.; Dang, T.; Nguyen, H.; Do, T.; Le, C.; Nguyen, H. T. T.; Le, N. Q.; Le, L. T.; Vu, D. M.; Ngo, T. D.; Le, H. T. T.; Nguyen, L. T.; Ha, T. C.; Hoang, Y.; Dao, D. X.; Giang, P. H.; Luu, H. N.; Dao, M. D.; Le, L.; Le, V. S.; Tran, T.; Nguyen, Q.; Le, D.-H.; Nguyen, D. T.; Vu, V. H.; Vo, N. S.
Show abstract
Vietnam, the 16th most populated nation, remains profoundly underrepresented in global genomic databases. Here, we present VN1K, a first-ever comprehensive and well-curated resource of multi-omics data with a wide-range of phenotypic information of 1,011 unrelated Vietnamese individuals. High-depth short-read whole-genome sequencing data were generated for all samples along with various - omic data, including microarray, long-read whole-genome sequencing, and RNA sequencing. Using a high-sensitivity variant detection pipeline, which included a pangenome graph reference and a deep-learning framework, we identified nearly 40 million variants of which 8.5 million are novel with nearly 900 thousand short insertions/deletions and 39 thousand structural variants. Specifically, VN1K featured a first-ever whole-genome methylation profile based on long read sequencing. A genotype imputation panel was also created with the highest accuracy on the Vietnamese population. Variants with significantly different allele frequencies in the Vietnamese population compared to others were found to be functionally significant, especially in genes associated with immune diseases (HLA-B, KIR3DL3, KIR2DL1, KIR2DL4) or drug responses (CYP2C19, CYP2D6, VKORC1, CYP2B6). We were also able to map various loci related to hepatitis B virus infection as well as six disease traits, including triglyceride levels, LDL-C, serum glucose levels, HbA1c, and levels of two liver enzymes (ALT and AST). VN1K dataset is accessible via genome.vinbigdata.org, an integrated platform with both linear and graph-based genome browser for facilitating data exploration, research, and applications in precision medicine.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Full resolution HLA and KIR genes annotation for human genome assemblies 95%
- Genetic Diversity and Structural Complexity of the Killer-Cell Immunoglobulin-Like Receptor Gene Complex: A Comprehensive Analysis using Human Pangenome Assemblies 94%
- Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation 94%
Similar papers in this journal
- Whole-genome reference panel of 1,781 Northeast Asians improves imputation accuracy of rare and low-frequency variants 94%
- Multi-modal investigation of the schizophrenia-associated 3q29 genomic interval reveals global genetic diversity with unique haplotypes and segments that increase the risk for non-allelic homologous recombination 94%
- Nanopore sequencing with unique molecular identifiers enables accurate mutation analysis and haplotyping in the complex Lipoprotein(a) KIV-2 VNTR 93%
Similar papers in this journal
- Quartet DNA reference materials and datasets for comprehensively evaluating germline variants calling performance 96%
- A Refined Analysis of Neanderthal-Introgressed Sequences in Modern Humans with a Complete Reference Genome 95%
- Characterization of Structural Variation in Tibetans Reveals New Evidence of High-altitude Adaptation and Introgression 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.