Back

Human Genetics

Springer Science and Business Media LLC

All preprints, ranked by how well they match Human Genetics's content profile, based on 28 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
The Genetic Landscape and Epidemiological Characteristics of Inherited Retinal Diseases in the Chinese Population

Zeng, B.; Cui, Z.; Zhou, S.; Dai, W.

2026-05-29 ophthalmology 10.64898/2026.05.27.26354224 medRxiv
Top 0.1%
33.7%
Show abstract

Background: Inherited Retinal Diseases (IRDs) are a group of genetically heterogeneous blinding conditions. Major global genomic reference databases are disproportionately enriched for individuals of European ancestry. This underrepresentation creates a significant bias that impedes the accuracy of genetic diagnosis in the Chinese population. This study aims to address this limitation by constructing a comprehensive genetic landscape of IRDs using large-scale deep-sequencing data from a large Chinese cohort. Methods: The study leveraged variant data primarily from 10,588 individuals in the China Metabolic Analytics Project (ChinaMAP) and cross-referenced findings against multiple national and international databases. We systematically curated variants within a targeted panel of 291 IRD-associated genes. Variant pathogenicity was assessed using a comprehensive pipeline integrating InterVar-automated classification based on 2015 American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines, ClinVar evidence (review status [≥] 1 star), and manual literature curation. We delineated the mutational spectrum, identified population-enriched pathogenic/likely pathogenic (P/LP) variants, and analyzed the distribution characteristics of IRD-associated highly-mutated genes. Furthermore, we calculated the carrier frequencies (CF) and genetic prevalence (GP) of autosomal recessive(AR)-IRD genes in the Chinese population. Results: The study revealed a highly concentrated genetic landscape for AR-IRDs in the Chinese population, with ABCA4 and USH2A emerging as the primary drivers of the genetic burden. This finding aligns with previous Chinese cohorts but contrasts with global databases, where genes such as the X-linked RPGR are more prevalent. In contrast, autosomal dominant (AD)-IRDs exhibited high locus heterogeneity, with pathogenic variants dispersed across numerous genes (e.g., COL2A1 and MFN2). We identified a series of P/LP variants that were either high-frequency or significantly enriched in the Chinese population, such as CNGB1 (p.P530R) and specific recurrent alleles in ABCA4 and CYP4V2. The estimated cumulative CF for AR-IRDs was 1 in 5.60, and the theoretical total GP was 1 in 2,624.67, based on the ChinaMAP data. Conclusion: By integrating the ChinaMAP dataset with diverse genomic resources, this study provides a genetic landscape of IRDs in the Chinese population. Our analysis shows a concentrated mutational spectrum in AR-IRDs, contrasting with the pronounced heterogeneity in AD-IRDs. These findings, including population-specific pathogenic variants and refined prevalence estimates, provide a resource for precision diagnostics, genetic counseling, expanded carrier screening (ECS), and public health policy development in China.

2
Genetic Evidence for Soluble VEGFR2 as a Protective Factor Against Macular Pucker

Gao, Q.; Guo, J.

2024-11-30 ophthalmology 10.1101/2024.11.29.24318222 medRxiv
Top 0.1%
18.7%
Show abstract

BackgroundIdentifying factors that protect against macular pucker (MP) is crucial for developing effective treatments. We hypothesize that soluble vascular endothelial growth factor receptor 2 (sVEGFR2), a decoy receptor for VEGF signaling, may provide such protection. MethodsWe performed a genome-wide association (GWAS) meta-analysis on three independent studies of sVEGFR2 protein quantitative trait loci (pQTL), which was used to identify novel variants and select variants strongly associated with sVEGFR2 levels. We investigated the association between sVEGFR2 levels and the risk of MP using cis-Mendelian Randomization (cis-MR) and Bayesian colocalization. We also used the UKB-PPP study of sVEGFR2 pQTL, which used a different method for measuring plasma protein, to validate the results from cis-MR and colocalization. ResultsGWAS meta-analysis identified four novel variants associated with sVEGFR2 levels, rs9411335 (MED27), rs6992534 (ERI1), rs11893824 (LIMS2), and rs7678559 (PCDH7). Cis-MR analysis suggested that higher levels of sVEGFR2, genetically predicted using the meta-analysis dataset, were associated with a lower risk of MP (OR 0.86, 95% CI 0.77-0.95, P = 4.6 x 10-3). Colocalization supported shared genetic variants in the VEGFR2 gene region between sVEGFR2 and MP (PPH4 = 0.92). Using sVEGFR2 pQTL from UKB-PPP study, sVEGFR2s inverse association with MP was validated by cis-MR (OR 0.86, 95% CI 0.78-0.96, P = 4.4 x 10-3) and colocalization (PPH4 = 0.95). Sensitivity analyses showed no evidence of pleiotropy or reverse causality, though some heterogeneity was found. ConclusionsOur study suggested that sVEGFR2 has a protective effect against MP and could be a potential therapeutic target. Further research is necessary to elucidate its protective mechanisms and validate its translational implications.

3
Rare Variants in the OTOG Gene Are a Frequent Cause of Familial Meniere’s Disease

Roman-Naranjo, P.; Gallego-Martinez, A.; Soto-Varela, A.; Aran, I.; Moleon, M. d. C.; Espinosa-Sanchez, J. M.; Amor-Dorado, J. C.; Batuecas-Caletrio, A.; Perez-Vazquez, P.; Lopez-Escamez, J. A.

2019-09-21 genetics 10.1101/771527 medRxiv
Top 0.1%
18.5%
Show abstract

ObjectivesMenieres disease (MD) is a rare inner ear disorder characterized by sensorineural hearing loss, episodic vertigo and tinnitus. Familial MD has been reported in 6-9% of sporadic cases, and few genes including FAM136A, DTNA, PRKCB, SEMA3D and DPT have been involved in single families, suggesting genetic heterogeneity. In this study, the authors recruited 46 families with MD to search for relevant candidate genes for hearing loss in familial MD.\n\nDesignExome sequencing data from MD patients were analyzed to search for rare variants in hearing loss genes in a case-control study. A total of 109 patients with MD (73 familial cases and 36 early-onset sporadic patients) diagnosed according to the diagnostic criteria defined by the Barany Society were recruited in 11 hospitals. The allelic frequencies of rare variants in hearing loss genes were calculated in individuals with familial MD. A single rare variant analysis (SRVA) and a gene burden analysis (GBA) were conducted in the dataset selecting one patient from each family. Allelic frequencies from European and Spanish reference datasets were used as controls.\n\nResultsA total of 5136 single nucleotide variants in hearing loss genes were considered for SRVA in familial MD cases, but only one heterozygous variant in the OTOG gene (rs552304627) was found in two unrelated families. The GBA found an enrichment of rare missense variants in the OTOG gene in familial MD. So, 15/46 families (33%) showed at least one rare missense variant in the OTOG gene, suggesting a key role in familial MD.\n\nConclusionsThe authors found an enrichment of multiplex rare missense variants in the OTOG gene in familial MD. This finding supports OTOG as a relevant gene in familial MD and set the groundwork for genetic testing in MD.

4
Genetic Dependence and Genetic Diseases

Li, B.; Bian, W.-J.; Zhou, P.; Wang, J.; Fan, C.-X.; Xu, H.-Q.; Yu, L.; He, N.; Shi, Y.-W.; Su, T.; Yi, Y.-H.; Liao, W.-P.

2023-08-05 genetics 10.1101/2023.08.02.551736 medRxiv
Top 0.1%
18.4%
Show abstract

The human life depends on the function of proteins that are encoded by about twenty-thousand genes. The gene-disease associations in majority genes are unknown and the mechanisms underlying pathogenicity of genes/variants and common diseases remain unclear. We studied how human life depends on the genes, i.e., the genetic-dependence, which was classified as genetic-dependent nature (GDN, vital consequence of abolishing a gene), genetic-dependent quantity (GDQ, quantitative genetic function required for normal life), and genetic-dependent stage (GDS, temporal expression pattern). Each gene differs in genetic-dependent features, which determines the gene-disease association extensively. The GDN is associated with the pathogenic potential/feature of genes and the strength of pathogenicity. The GDQ-damage relation determines the pathogenicity of variants and subsequently the pathogenic genotype, phenotype spectrum, and inheritance of variants. The GDS is mainly associated with the onset age/evolution/outcome and the nature of genetic disorders (disease/susceptibility). The varied and quantitative genetic-dependent feature of genome explains common mild phenotype/susceptibility. The genetic-dependence discloses the mechanisms underlying pathogenicity of gene/variants and common diseases. One sentence summaryGenetic dependent feature differs in genes and determines pathogenicity of genes/variants and the clinical features of genetic diseases.

5
The prevalence of protein misfolding as a mechanism for hereditary deafness

Gogal, R. A.; Cox, G. M.; Kolbe, D. L.; Odell, A. M.; Ovel, C. E.; McCormick, K. I.; Hong, B.; Azaiez, H.; Casavant, T. L.; Smith, R. J. H.; Braun, T. A.; Schnieders, M. J.

2026-03-11 genetics 10.64898/2026.03.09.710547 medRxiv
Top 0.1%
18.3%
Show abstract

Hearing loss is the most common sensory deficit impacting [~]5% of the worlds population. The Deafness Variation Database (DVD) is a public resource of deafness variants, containing over 380,000 missense variants across 224 genes, with 303,577 classified as a variant of uncertain significance (VUS). To address the challenge of evaluating each deafness associated VUS, we evaluate a family of probabilistic frameworks to quantify the strength of computational evidence based on ACMG/AMP recommendations. First, CADD and REVEL are compared using Bayesian models parameterized using either a ClinVar 2019 dataset or labeled DVD variants. The REVEL model built using the DVD dataset demonstrates the best accuracy, sensitivity, and specificity. Incorporation of (in)tolerance to missense variation based on sorting each gene into three bins (tolerant, average, intolerant) shows that intolerant DVD genes are consistent with a higher prior probability of being pathogenic (25.7%) than average (10.7%) or tolerant (8.7%) genes. Finally, the impact of protein folding stability was incorporated using a 2D likelihood, which surpassed the simpler models while also offering a biophysical rationale for the disease mechanism. The protein folding-informed Bayesian model results in 28,866 prioritized VUSs reaching a posterior probability of pathogenicity above 98% with a false positive rate of only 0.14%. Overall, 54,752 missense variants are predicted to cause protein folding destabilization of greater than 1.0 kcal/mol, while 18,706 of the 28,886 prioritized VUS (65%) surpass this threshold. From these VUSs, we identify twelve probands where the patients genetic diagnosis is upgraded to likely pathogenic/pathogenic. We highlight two variants that cause clear structural disruption, demonstrating the impact of biophysical characterization on variant evaluation. Author SummaryWe investigate the impacts of single amino acid changes on protein structure and folding in the context of hearing loss. Hearing loss is the most common impairment of the main senses affecting nearly 5% of the worlds population. About 45% of people with hearing loss receive a diagnosis after targeted genetic testing. Here, we integrate biophysical data that quantifies the effect of a change to protein sequence on protein folding in combination with genetic data to improve our ability to identify protein amino acid changes that are likely to impact hearing. Our work leads to 12 patients receiving an upgraded diagnosis with their variant disrupting protein stability. Although the method is applied to hearing loss, it can be used for interpreting protein sequence changes in other disease contexts.

6
Replication of missense OTOG gene variants in a Brazilian cohort of Meniere's Disease

Bianco-Bortoletto, G.; Almeida Carneiro, G.; Fabbri-Scallet, H.; Parra-Perez, A. M.; de Carvalho Lopes, K.; de Almeida Lima Sa Vieria, T.; Freitas Gananca, F.; Amor-Dorado, J. C.; Soto-Valera, A.; Lopez-Escamez, J. A.; Sartorato, E. L.

2025-04-28 genetic and genomic medicine 10.1101/2025.04.26.25326273 medRxiv
Top 0.1%
15.1%
Show abstract

Menieres Disease (MD) is a chronic inner ear disorder defined by recurring episodes of vertigo, fluctuating sensorineural hearing loss, tinnitus, and/or fullness in the ear. Its prevalence varies by region and ethnicity, with scarce epidemiological data in Brazilian population. Although most MD cases are sporadic, Familial MD (FMD) is observed in 5% to 20% of European cases. By exome sequencing, we have found a rare missense variant in the OTOG gene in a Brazilian MD individual with probable European ancestry (chr11:17599671C>T), which was previously reported in a Spanish cohort. Two additional rare missense heterozygous OTOG variants were found in the same proband. Splice Site analysis showed that chr11:17599671C>T may lead to substantial changes generating exonic cis regulatory elements, and protein modelling revealed structural changes in the presence of chr11:17599671C>T, chr11:17576581G>C and chr11:17594108C>T, predicted to highly destabilize protein structure. These findings indicate that missense variants may have an additive effect leading to an unstable Otogelin and support OTOG gene as a key player in the MD pathophysiology.

7
A novel PS4 criterion approach based on symptoms of rare diseases and in-house frequency data in a Bayesian framework.

Cho, Y. K.; Won, D.-g.; Keum, C.; Seo, G. H.; Lee, B. H.; Lee, B.-C.

2020-07-24 genetics 10.1101/2020.07.22.215426 medRxiv
Top 0.1%
15.1%
Show abstract

The American College of Medical Genetics (ACMG) and Genomics/Association for Molecular Pathology (AMP) previously reported standardized guidance for the assessment of genetic variants. One of the criteria regarding the prevalence in a case-control study, PS4, is important due to its evidence of pathogenicity. Despite recent studies approaching gene- and disease-specific probands, interpretation of a variant to PS4 still has certain limitations for rare variants. Here, we suggest a generalized method, Bayesian odds ratio (BayesianOR), applicable to PS4 via decomposing a disease to its symptoms and applying a Bayesian framework. Using this approach, we demonstrate reproducibility of the calculation of the original odds ratio from well-studied epilepsy data and verify the applicability to in-house frequencies for various rare diseases. In addition, BayesianOR showed a significant difference in tendency with different ClinVar pathogenicity, using in-house data. Thus, the novel method described here should provide an improved interpretation of sequence variants. Furthermore, we anticipate that it will enhance the diagnosis of patients with rare diseases.

8
Regulatory Variants on the Leukocyte Immunoglobulin-Like Receptor Gene Cluster are Associated with Crohn's Disease and Interact with Regulatory Variants for TAP2

Kim, K.; Oh, S. J.; Lee, J.; Kwon, A.; Yu, C.-Y.; Kim, S.; Choi, C. H.; Kang, S.-B.; Kim, T. O.; Park, D. I.; Lee, C. K.

2023-03-28 genetic and genomic medicine 10.1101/2023.03.28.23287842 medRxiv
Top 0.1%
13.2%
Show abstract

Background and AimsCrohns disease (CD) has a complex polygenic etiology with high heritability. We keep putting an effort to identify novel variants associated with susceptibility to CD through a genome-wide association study (GWAS) in large Korean populations. MethodsGenome-wide variant data from 902 Korean patients with CD and 72,179 controls were used to assess the genetic associations in a meta-analysis with previous Korean GWAS results from 1,621 patients with CD and 4,419 controls. Epistatic interactions between CD-risk variants of interest were tested using a multivariate logistic regression model with an interaction term. ResultsWe identified two novel genetic associations with the risk of CD near ZBTB38 and within the leukocyte immunoglobulin-like receptor (LILR) gene cluster (P<5x10-8), with highly consistent effect sizes between the two independent Korean cohorts. CD-risk variants in the LILR locus are known quantitative trait loci (QTL) for multiple LILR genes, of which LILRB2 directly interacts with various ligands including MHC class I molecules. The LILR lead variant exhibited a significant epistatic interaction with CD-associated regulatory variants for TAP2 involved in the antigen presentation of MHC class I molecules (P=4.11x10-4), showing higher CD-risk effects of the TAP2 variant in individuals carrying more risk alleles of the LILR lead variant (OR=0.941, P=0.686 in non-carriers; OR=1.45, P=2.51x10-4 in single-copy carriers; OR=2.38, P=2.76x10-6 in two-copy carriers). ConclusionsThis study demonstrated that genetic variants at two novel susceptibility loci and the epistatic interaction between variants in LILR and TAP2 loci confer risk of CD.

9
Integrating genome-wide association and transcriptome predicted model identify novel target genes with osteoporosis

Yin, P.; Zhu, M.

2019-09-16 genetics 10.1101/771543 medRxiv
Top 0.1%
12.9%
Show abstract

Osteoporosis (OP) is a highly polygenetic disease which is usually characterized by low bone mineral density. Genome-wide association studies (GWAS) have identified hundreds of genetic loci associated with bone mineral density. However, the biological mechanisms of these loci remain elusive. To identify potential causal genes of the associated loci, we detected trait-gene expression associations by transcriptome-wide association study (TWAS) method. It directly imputes gene expression effects from GWAS data, using a statistical prediction model trained on GTEx reference transcriptome data, with blood and skeletal tissues data. Then we performed a colocalization analysis to evaluate the posterior probability of biological patterns: association characterized by a single shared causal variant or two distinct causal variants. The ultimate analysis identified 276 candidate genes, including 3 novel loci, 204 novel candidate genes and 69 replicated from GWAS. The 3 novel loci located at chr6: 72417543, chr15: 69601206, chr21: 30530692, mapping to gene RIMS1, SPESP1, MAP3K7CL. The results of colocalization analysis indicated that 142 of them showing strong evidence of a single shared causal variant and 134 of them showing evidence of joint causal variants. Their biological function was directly or indirectly associated with the occurrence of OP validated by VarElect tool. Several important OP-associated pathways were detected by protein-protein interaction and pathway enrichment analysis. Target genes were further enriched for differential expression genes in osteoblasts expression profiles, e.g. IBSP, affecting calcium and hydroxyapatite binding, and CD44, regulating alternative splicing of gene transcription. Transcriptome fine-mapping identifies more disease-related genes and provide additional insight into the development of novel targeted therapeutics to treat OP.

10
Spatial Distribution of Missense Variants within Complement Proteins Associates with Age Related Macular Degeneration

Grunin, M.; de Jong, S.; Palmer, E. L.; Jin, B.; Rinker, D.; Moth, C.; Capra, J. A.; Haines, J. L.; Bush, W.; den Hollander, A.; International Age-related Macular Degeneration Genomics Consortium,

2023-08-31 genetic and genomic medicine 10.1101/2023.08.28.23294686 medRxiv
Top 0.1%
12.4%
Show abstract

PurposeGenetic variants in complement genes are associated with age-related macular degeneration (AMD). However, many rare variants have been identified in these genes, but have an unknown significance, and their impact on protein function and structure is still unknown. We set out to address this issue by evaluating the spatial placement and impact on protein structureof these variants by developing an analytical pipeline and applying it to the International AMD Genomics Consortium (IAMDGC) dataset (16,144 AMD cases, 17,832 controls). MethodsThe IAMDGC dataset was imputed using the Haplotype Reference Consortium (HRC), leading to an improvement of over 30% more imputed variants, over the original 1000 Genomes imputation. Variants were extracted for the CFH, CFI, CFB, C9, and C3 genes, and filtered for missense variants in solved protein structures. We evaluated these variants as to their placement in the three-dimensional structure of the protein (i.e. spatial proximity in the protein), as well as AMD association. We applied several pipelines to a) calculate spatial proximity to known AMD variants versus gnomAD variants, b) assess a variants likelihood of causing protein destabilization via calculation of predicted free energy change (ddG) using Rosetta, and c) whole gene-based testing to test for statistical associations. Gene-based testing using seqMeta was performed using a) all variants b) variants near known AMD variants or c) with a ddG >|2|. Further, we applied a structural kernel adaptation of SKAT testing (POKEMON) to confirm the association of spatial distributions of missense variants to AMD. Finally, we used logistic regression on known AMD variants in CFI to identify variants leading to >50% reduction in protein expression from known AMD patient carriers of CFI variants compared to wild type (as determined by in vitro experiments) to determine the pipelines robustness in identifying AMD-relevant variants. These results were compared to functional impact scores, ie CADD values > 10, which indicate if a variant may have a large functional impact genomewide, to determine if our metrics have better discriminative power than existing variant assessment methods. Once our pipeline had been validated, we then performed a priori selection of variants using this pipeline methodology, and tested AMD patient cell lines that carried those selected variants from the EUGENDA cohort (n=34). We investigated complement pathway protein expression in vitro, looking at multiple components of the complement factor pathway in patient carriers of bioinformatically identified variants. ResultsMultiple variants were found with a ddG>|2| in each complement gene investigated. Gene-based tests using known and novel missense variants identified significant associations of the C3, C9, CFB, and CFH genes with AMD risk after controlling for age and sex (P=3.22x10-5;7.58x10-6;2.1x10-3;1.2x10-31). ddG filtering and SKAT-O tests indicate that missense variants that are predicted to destabilize the protein, in both CFI and CFH, are associated with AMD (P=CFH:0.05, CFI:0.01, threshold of 0.05 significance). Our structural kernel approach identified spatial associations for AMD risk within the protein structures for C3, C9, CFB, CFH, and CFI at a nominal p-value of 0.05. Both ddG and CADD scores were predictive of reduced CFI protein expression, with ROC curve analyses indicating ddG is a better predictor (AUCs of 0.76 and 0.69, respectively). A priori in vitro analysis of variants in all complement factor genes indicated that several variants identified via bioinformatics programs PathProx/POKEMON in our pipeline via in vitro experiments caused significant change in complement protein expression (P=0.04) in actual patient carriers of those variants, via ELISA testing of proteins in the complement factor pathway, and were previously unknown to contribute to AMD pathogenesis. ConclusionWe demonstrate for the first time that missense variants in complement genes cluster together spatially and are associated with AMD case/control status. Using this method, we can identify CFI and CFH variants of previously unknown significance that are predicted to destabilize the proteins. These variants, both in and outside spatial clusters, can predict in-vitro tested CFI protein expression changes, and we hypothesize the same is true for CFH. A priori identification of variants that impact gene expression allow for classification for previously classified as VUS. Further investigation is needed to validate the models for additional variants and to be applied to all AMD-associated genes.

11
Effects of Vascular Endothelial Growth Factor Family on Macular Pucker: cis-Mendelian Randomization and Colocalization Analyses

Gao, Q.; Guo, J.

2024-12-01 ophthalmology 10.1101/2024.12.01.24318223 medRxiv
Top 0.1%
12.2%
Show abstract

BackgroundVascular endothelial growth factor (VEGF) family and its receptors (VEGFR) could be implicated in macular pucker (MP) pathogenesis. We used cis-Mendelian randomization (cis-MR) and Bayesian colocalization with summary-level genome-wide association study (GWAS) data to explore causal relationships. MethodsGenetic variants associated with soluble VEGFR2 (sVEGFR2), sVEGFR3, VEGF-A, VEGF-C, or VEGF-D levels were selected from a GWAS of protein quantitative trait loci with 35,559 Icelanders. MP GWAS (3,974 cases, 376,650 controls) was sourced from the FinnGen. We employed cis-MR using invariance-weighted median, supplemented by other methods. Bayesian colocalization validated cis-MR findings. Pleiotropy, reverse causality, and heterogeneity were assessed. FindingsCis-MR suggested that genetically predicted higher levels of sVEGFR2 were associated with reduced MP risk (Odds ratio (OR) 0.82, 95% confidence interval (CI) 0.75-0.89, P=9.20 x 10-6). Colocalization supported shared genetic variants in the VEGFR2 gene region between sVEGFR2 and MP (posterior probability of hypothesis 4 (PPH4) =0.94), reinforcing sVEGFR2s protective role in MP. Although cis-MR suggested an inverse relationship between sVEGFR3 levels and MP risk (OR 0.71, 95% CI 0.57-0.89, P=2.64 x 10-3), colocalization analysis did not confirm direct causality (PPH4=0.03). VEGF-A, VEGF-C and VEGF-D levels were not associated with MP risk in MR analyses. No evidence of pleiotropy, reverse causality, and heterogeneity was found across our MR analyses. InterpretationOur study suggested that sVEGFR2 has causally protective effects against MP and could serve as a potential drug target. Further research is necessary to elucidate its protective mechanisms and validate its translational implications.

12
Disease-causing variant recommendation system for clinical genome interpretation with adjusted scores for artefactual variants

Kim, H. H.; Woo, J.; Kim, D.-W.; Lee, J.; Seo, G. H.; Lee, H.; Lee, K.

2022-10-14 genetics 10.1101/2022.10.12.511857 medRxiv
Top 0.1%
12.2%
Show abstract

BackgroundIn the process of finding the causative variant of rare diseases (RD), accurate assessment and prioritization of genetic variants is essential. Although quality control (QC) of genetic variants is strictly performed, the presence of artefactual variants in the remaining set of variants can deteriorate the process. Variant QC and prioritization have been treated as separate processes, leading to limited efficiency and risk of misdiagnosis. ResultsWe developed a disease-causing variant recommendation system that integrates quality control into variant prioritization by adjusting scores for artefactual variants. We confirmed that the QC-related features of the variants contribute to a significant performance improvement. For genomic data from 2,878 patients with rare disorders, the recall rate of finding causative variants was 0.961 for the top 5 ranked variants. We also found that our system recognized the anomaly of QC-related features, so that the scores of artifactual variants to be disease-causing were assessed relatively low. ConclusionsIntegration of variant QC and prioritization help reduce the risk of misdiagnosis based on artefactual variants and increase the effectiveness of clinical genome interpretation.

13
COVID-19 relevant genetic variants confirmed in an admixed population

Texis, T.; Cruz-Jaramilllo, J. L.; Garcia-Munoz, W.; Anzures-Cortes, L.; Hadadd-Talancon, L.; Sanchez-Garcia, S.; Jimenez-Martinez, M. d. C.; Perez-Barragan, E.; Nieto-Patlan, A.; Martinez-Ezquerro, J. D.; Rubio-Carrasco, K.; Rodriguez-Dorantes, M.; Cortes-Ramirez, S.; Mellado-Sanchez, G.; Perez-Tapia, S. M.; Gonzalez-Covarrubias, V.

2022-04-16 genetic and genomic medicine 10.1101/2022.04.15.22273925 medRxiv
Top 0.1%
11.9%
Show abstract

The dissection of factors that contribute to COVID-19 infection and severity has overwhelmed the scientific community for almost 2 years. Current reports highlight the role of in disease incidence, progression, and severity. Here, we aimed to confirm the presence of previously reported genetic variants in an admixed population. Allele frequencies were assessed and compared between the general population (N=3079) for which at least 30% have not been infected with SARS-CoV2 as per July 2021 versus COVID-19 patients (N=106). Genotyping data from the Illumina GSA array was used to impute genetic variation for 14 COVID-relevant genes, using the 1000G phase 3 as reference based on the human genome assembly hg19, following current standard protocols and recommendations for genetic imputation. Bioinformatic and statistical analyses were performed using MACH v1.0, R, and PLINK. A total of 7953 variants were imputed on, ABO, CCR2, CCR9, CXCR6, DPP9, FYCO1, IL10RB/IFNAR2, LZTFL1, OAS1, OAS2, OAS3, SLC6A20, TYK2, and XCR1. Statistically significant allele differences were reported for 10 and 7 previously identified and confirmed variants, ABO rs657152, DPP9 rs2109069, LZTFL1 rs11385942, OAS1 rs10774671, OAS1 rs2660, OAS2 rs1293767, and OAS3 rs1859330 p<0.03. In addition, we identified 842 variants in these COVID-related genes with significant allele frequency differences between COVID patients and the general population (p-value <E-2 - E-179). Our observations confirm the presence of genetic differences in COVID-19 patients in an admixed population and prompts for the investigation of the statistical relevance of additional variants on these and other genes that could identify local and geographical patterns of COVID-19.

14
A Foundational Exome Resource for Jordan: Dual Ancestry Admixture and Population-Specific Variants to Improve Clinical Variant Interpretation

Froukh, T.

2026-05-27 genetic and genomic medicine 10.64898/2026.05.23.26353895 medRxiv
Top 0.1%
11.7%
Show abstract

Currently, the genetic architecture of Middle Eastern populations is underrepresented in global genomic databases. This gap increases the rate of Variants of Uncertain Significance (VUSs) and clinical misinterpretations of genomic data especially in Middle Eastern populations. Whole exome sequencing was conducted on 90 healthy individuals from Jordan and the data were analysed using Principal Component Analysis (PCA) and multi-computational filtering. PCA revealed a double ancestry (EUR-AFR) admixture rather than a triple admixture (EUR-AFR-AMR). More than 3,500 populations-specific variants (PSVs) were identified, of which 72% were singletons. Additionally, 19 variants were significantly enriched compared to the maximum allele frequencies in public global databases (Fisher's exact test with Benjamini-Hochberg false discovery rate correction, p-value < 0.05). Consequently, the results suggest the reclassification of variants of Uncertain Significance (VUS) which reside in the ECE2 gene to likely benign and the variants of Conflicting Classification of Pathogenicity in the genes IL1RN and THPO to benign based on the significant allele frequency (AF=0.0389, p-value < 0.05). Furthermore, a pathogenic ClinVar variant was identified in a healthy individual, warranting careful interpretation. The findings underscore the importance of identifying PSVs in order to minimize or even prevent clinical misdiagnosis and highlight the unique genetic signature in Jordan. The study serves as a foundational resource for precision medicine in the region.

15
Complete genomic profiles of 1,496 Taiwanese reveal curated medical insights

Wu, D.-C.; Hsu, J. S.; Chen, C.-Y.; Shih, S.-H.; Liu, J.-F.; Tsai, Y.-C.; Lee, T.-L.; Chen, W.-A.; Tseng, Y.-H.; Lo, Y.-C.; Lin, H.-Y.; Chen, Y.-C.; Chen, J.-Y.; Chang, T.-H.; Guo, W.-H.; Mao, H.-H.; Chen, P.-L.

2021-12-30 genetic and genomic medicine 10.1101/2021.12.23.21268291 medRxiv
Top 0.1%
11.7%
Show abstract

BackgroundTaiwan Biobank (TWB) project has built a nationwide database to facilitate the basic and clinical collaboration within the island and internationally, which is one of the valuable public datasets of the East Asian population. This study provided comprehensive genomic medicine findings from 1,496 WGS data from TWB. MethodsWe reanalyzed 1,496 Illumina-based whole genome sequences (WGS) of Taiwanese participants with at least 30X depth of coverage by Sentieon DNAscope, a precisionFDA challenge winner method. All single nucleotide variants (SNV) and small insertions/deletions (Indel) have been jointly called and recalibrated as one cohort dataset. Multiple practicing clinicians have reviewed clinically significant variants. ResultsWe found that each Taiwanese has 6,870.7 globally novel variants and classified all genomic positions according to the recalibrated sequence qualities. The variant quality score helps distinguish actual genetic variants among the technical false-positive variants, making the accurate variant minor allele frequency (MAF). All variant annotation information can be browsed at TaiwanGenomes (https://genomes.tw). We detected 54 PharmGKB-reported Cytochrome P450 (CYP) genes haplotype-drug pairs with MAF over 10% in the TWB cohort and 39.8% (439/1103) Taiwanese harbored at least one PharmGKB-reported human leukocyte antigen (HLA) risk allele. We also identified 23 variants located at ACMG secondary finding V3 gene list from 25 participants, indicating 1.67% of the population is harboring at least one medical actionable variant. For carrier status of all known pathogenic variants, we estimated one in 22 couples (4.52%) would be under the risk of having offspring with at least one pathogenic variant, which is in line with Japanese (JPN) and Singaporean (SGN) populations. We also detected 6.88% and 2.02% of carrier rates for alpha thalassemia and spinal muscular atrophy (SMA) for copy number pathogenic variants, respectively. ConclusionAs WGS has become affordable for everyone, a person only needs to test once for a lifetime; comprehensive WGS data reanalysis of the genomic profile will have a significant clinical impact. Our study highlights the overall picture of a complete genomic profile with medical information for a population and individuals.

16
COVID-19 risk haplogroups differ between populations, deviate from Neanderthal haplotypes and compromise risk assessment in non-Europeans

Wohlers, I.; Calonga-Solis, V.; Jobst, J.-N.; Busch, H.

2020-11-03 genetics 10.1101/2020.11.02.365551 medRxiv
Top 0.1%
11.7%
Show abstract

Recent genome wide association studies (GWAS) have identified genetic risk factors for developing severe COVID-19 symptoms. The first published study reported a 1bp insertion rs11385942 on chromosome 3 (1) and subsequent studies single nucleotide variants (SNVs) such as rs35044562, rs67959919 (2) and rs13078854 (3), all highly correlated with each other. Zeberg and Paabo (4) subsequently traced them back to Neanderthal origin. They found that a 49.4 kb genomic region including the risk allele of rs35044562 is inherited from Neanderthals of Vindija in Croatia. Here we add a differently focused evaluation of this major genetic risk factor to these recent analyses. We show that (i) COVID-19-related genetic factors of three previously assessed Neanderthals deviate from those of modern humans and that (ii) they differ among world-wide human populations, which compromises risk prediction in non-Europeans. Currently, caution is thus advised in the genetic risk assessment of non-Europeans during this world-wide COVID-19 pandemic.

17
Whole exome sequencing study identifies candidate loss of function variants and locus heterogeneity in familial cholesteatoma

Cardenas, R. P.; Prinsley, P.; Philpott, C.; Bhutta, M.; Wilson, E.; Brewer, D.; Jennings, B.

2022-07-16 genetics 10.1101/2022.07.15.500191 medRxiv
Top 0.1%
11.0%
Show abstract

Cholesteatoma is a rare progressive disease of the middle ear. Most cases are sporadic, but some patients report a positive family history. Identifying functionally important gene variants associated with this disease has the potential to uncover the molecular basis of cholesteatoma pathology with implications for disease prevention, surveillance, or management. We performed an observational WES study of 21 individuals treated for cholesteatoma who were recruited from ten multiply affected families. These family studies were complemented with gene-level mutational burden analysis. We also applied functional enrichment analyses to identify shared properties and pathways for candidate genes and their products. Filtered data collected from pairs and trios of participants within the ten families revealed 398 rare, loss of function (LOF) variants co-segregating with cholesteatoma in 389 genes. We identified six genes DENND2C, DNAH7, NBEAL1, NEB, PRRC2C, and SHC2, for which we found LOF variants in two or more families. The parallel gene-level of mutation-burden identified a significant mutation burden for the genes in the DNAH gene family, which encode products involved in ciliary structure. Functional enrichment analyses identified common pathways for the candidate genes which included GTPase regulator activity, calcium ion binding, and degradation of the extracellular matrix. The number of candidate genes identified and the locus heterogeneity that we describe within and between multiply affected families suggest that the genetic architecture for familial cholesteatoma is complex.

18
RGnet: Recessive Genotype Network in a Large Mendelian Disease Cohort

Ai, F.; Kang, L.; Zeng, J.; He, M.; Zhong, M.; Cheng, J.; Lu, Y.; Yuan, H.; Bu, F.

2024-12-04 genetic and genomic medicine 10.1101/2024.12.02.24318353 medRxiv
Top 0.1%
10.6%
Show abstract

Recessive genotypes, including compound heterozygotes and homozygotes formed by rare variants that impact gene function, affect both alleles and were linked to numerous diseases and traits. However, the underlying patterns and interconnections of these recessive genotypes in large cohorts have rarely been studied. To address this gap, the Recessive Genotype Network (RGnet) was developed. This network model maps variant and genotype features to visualize and analyze recessive genotype patterns within large cohorts. Additionally, it uses permutation-based analyses to assess the enrichment of these genotypes in relation to specific phenotypes. Demonstrated through its application to the genetic deafness gene SLC26A4 in 22,125 cases affected by hearing loss, RGnet successfully identified pathogenic variants with high connectivity, providing a reliable method for exploring the pathogenic mechanisms underlying recessive disorders or traits. Availability and ImplementationRGnet is available from GitHub at https://github.com/jiayiiiZeng/RGnet Contactbufengxiao@wchscu.cn Supplementary informationSupplementary data are available at Bioinformatics online.

19
Refining the genetic landscape of anophthalmia and microphthalmia: a comprehensive framework with deep learning and updated gene panels

Maftei, M. I.; Spink, L. G. N.; Carmona, O. G.; Mrstakova, S. M.; Abahreh, L.; Hayes, R.; Banon, A.; Cuevas, M. E.; Cid, K.; Araya-Secchi, R.; Fraternali, F.; Yu, J.; Arno, G.; Young, R.

2025-08-28 ophthalmology 10.1101/2025.08.26.25334245 medRxiv
Top 0.1%
9.9%
Show abstract

ImportanceAnophthalmia and microphthalmia (A/M) are rare congenital eye disorders with a low molecular diagnosis rate, which limits clinical management and genetic counselling. Improved detection and interpretation of pathogenic variants is essential for advancing diagnosis and care in affected individuals. ObjectiveTo improve the molecular diagnostic yield in A/M patients by refining the methodology of variant investigation and association using an updated rigorously curated gene panel, and a refined bioinformatic pipeline incorporating structural variant detection, in silico Artificial Intelligence assisted predictive tools, and molecular dynamics simulations. MethodologyWe curated an updated A/M gene panel through a systematic literature review and screened for rare variants in these genes using data from the UKs 100,000 Genomes Project, a national whole-genome sequencing initiative conducted by Genomics England. The cohort comprised 306 individuals recruited to the Rare Disease programme with a clinical diagnosis of anophthalmia or microphthalmia, recorded either as the primary phenotype or within HPO, SNOMED, or ICD-10 terms. Variants, including loss-of-function, missense, RNA splicing, and structural variants, were annotated with deep learning tools (AlphaMissense, SpliceAI), and missense variants were further assessed using REVEL, Missense3D, and molecular dynamics simulations. ResultsWe identified pathogenic or likely pathogenic variants in 37 (12.1%) individuals, with an additional 23 (7.5%) harbouring strong candidate variants of uncertain significance. Our literature review identified the biggest contributors to A/M phenotypes to be MFRP, OTX2, PRSS56 and SOX2, each with over 100 patients reported in the literature, with a total number of 124 genes found be associated to A/M. Variants from our screen were most often found in genes with high A/M association, but also included novel findings within genes with a weaker association to A/M such as ACTG1, HDAC6, RERE and SIX3; adding support to their disease relevance. Conclusions and RelevanceThis study increased the diagnostic yield in A/M patients recruited to the 100KGP, and provides further evidence of genotype-phenotype associations within the aetiology of A/M. We also provide an updated framework for enhancing clinical genetic diagnosis in A/M that may inform broader strategies for other complex congenital disorders. However, as molecular diagnosis of A/M remains low, further research in understanding the genetic aetiology of A/M is necessary.

20
Resolving the dark matter of ABCA4 for 1,054 Stargardt disease probands through integrated genomics and transcriptomics

Khan, M.; Cornelis, S. S.; Pozo-Valero, M. d.; Whelan, L.; Runhart, E. H.; Mishra, K.; Bults, F.; AlSwaiti, Y.; AlTabishi, A.; Baere, E. D.; Banfi, S.; Banin, E.; Bauwens, M.; Ben-Yosef, T.; Boon, C. J. F.; Born, L. I. v. d.; Defoort, S.; Devos, A.; Dockery, A.; Dudakova, L.; Fakin, A.; Farrar, G. J.; Ferraz Sallum, J. M.; Fujinami, K.; Gilissen, C.; Glavac, D.; Gorin, M. B.; Greenberg, J.; Hayashi, T.; Hettinga, Y.; Hoischen, A.; Hoyng, C. B.; Hufendiek, K.; Jagle, H.; Kamakari, S.; Karali, M.; Kellner, U.; Klaver, C. C. W.; Kousal, B.; Lamey, T.; MacDonald, I. M.; Matynia, A.; McLaren, T.; M

2019-10-25 genetics 10.1101/817767 medRxiv
Top 0.1%
9.8%
Show abstract

Missing heritability in human diseases represents a major challenge. Although whole-genome sequencing enables the analysis of coding and non-coding sequences, substantial costs and data storage requirements hamper its large-scale use to (re)sequence genes in genetically unsolved cases. The ABCA4 gene implicated in Stargardt disease (STGD1) has been studied extensively for 22 years, but thousands of cases remained unsolved. Therefore, single molecule molecular inversion probes were designed that enabled an automated and cost-effective sequence analysis of the complete 128-kb ABCA4 gene. Analysis of 1,054 unsolved STGD and STGD-like probands resulted in bi-allelic variations in 448 probands. Twenty-seven different causal deep-intronic variants were identified in 117 alleles. Based on in vitro splice assays, the 13 novel causal deep-intronic variants were found to result in pseudo-exon (PE) insertions (n=10) or exon elongations (n=3). Intriguingly, intron 13 variants c.1938-621G>A and c.1938-514G>A resulted in dual PE insertions consisting of the same upstream, but different downstream PEs. The intron 44 variant c.6148-84A>T resulted in two PE insertions that were accompanied by flanking exon deletions. Structural variant analysis revealed 11 distinct deletions, two of which contained small inverted segments. Uniparental isodisomy of chromosome 1 was identified in one proband. Integrated complete gene sequencing combined with transcript analysis, identified pathogenic deep-intronic and structural variants in 26% of bi-allelic cases not solved previously by sequencing of coding regions. This strategy serves as a model study that can be applied to other inherited diseases in which only one or a few genes are involved in the majority of cases.