Complex structural variation, phylogeny, and disease associations of the mucin pangenome
Plender, E. G.; Prodanov, T.; Lin, J.; Wong, I.; Wertz, J.; Gordon, W. W.; Bamshad, M. J.; Munson, K. M.; O'Neal, W. K.; Bloom, J. D.; Human Pangenome Reference Consortium, ; Marschall, T.; Eichler, E. E.
Show abstract
Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving [≥]97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range ({Delta} = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves [≥]95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- TAD Evolutionary and functional characterization reveals diversity in mammalian TAD boundary properties and function 97%
- A cell type-aware framework for nominating non-coding variants in Mendelian regulatory disorders 97%
- Long-read transcriptomics of a diverse human cohort reveals widespread ancestry bias in gene annotation 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.