Automatic variant prioritization in suspected genetic kidney disease using the Nephro Candidate Score (N-CS)
Rank, N.; Lukassen, S.; Anderegg, M.; Eckardt, K.-U.; Halbritter, J. P.; Popp, B.
Show abstract
Research QuestionDespite the identification of >700 genes linked to rare and inherited kidney diseases (IKD), many individuals with presumed IKD do not receive a diagnosis through genetic testing of known disease genes. Therefore, the identification of new disease genes is crucial to ending diagnostic odysseys, improving genetic counseling, and expanding treatment options. While the generation of large-scale sequencing data is no longer a substantial bottleneck, its interpretation remains challenging and offers room for improvement, notably in the discovery of novel disease genes. MethodsWe developed the Nephro Candidate Score (N-CS), a machine learning (ML) tool that prioritizes variants by combining a Nephro Gene Score (N-GS), a Nephro Variant Score (N-VS), and an Inheritance Score (IS). The ML-based N-GS and N-VS were trained on a wide range of genomic features to predict gene-disease relevance and variant pathogenicity, while the IS incorporates the mode of inheritance via a scoring heuristic. A Gene Set Enrichment Analysis (GSEA) was used to test whether genes top-ranked by the N-GS were enriched for kidney-related biological processes. Additionally, we tested the N-CS on an independent set of novel IKD candidate genes identified through a systematic literature search to validate its real-world performance. ResultsThe machine learning models for the N-CS subscores demonstrated high predictive accuracy, with an XGBoost algorithm for the N-GS achieving an AUC of 0.94 and a Logistic Regression model for the N-VS reaching an AUC of 0.99 in independent test sets. The biological relevance of the N-GS ranking was confirmed by the GSEA showing a significant enrichment of kidney-associated biological processes among top-scoring genes (p < 0.001). In the independent validation using recently published literature, the N-CS assigned compellingly high scores to the majority (10 of 11) of novel candidate genes for kidney disease, demonstrating its ability to generalize to new discoveries. ConclusionThe N-CS is a robust digital solution that can accelerate disease gene discovery and comes with the potential to reduce time to diagnosis. To support standardization and collaboration, the full N-CS framework is freely available, including a user-friendly web tool (NC-Scorer: https://nc-scorer.kidney-genetics.org/) and a command-line interface for high-throughput analysis, enabling standardized, sharable evaluation of candidate variants.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Single-cell RNA sequencing reveals mRNA splice isoform switching during kidney development 94%
- Computational characterization of lymphocyte topology on whole slide images of glomerular diseases 93%
- Multi-trait Analysis of GWAS for circulating FGF23 Identifies Novel Network Interactions Between HRG-HMGB1 and Cardiac Disease in CKD 93%
Similar papers in this journal
- Latent disease similarities and therapeutic repurposing possibilities uncovered by multi-modal generative topic modeling of human diseases 93%
- VAREANT : a bioinformatics application for gene variant reduction and annotation 91%
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 90%
Similar papers in this journal
- A functional landscape of chronic kidney disease entities from public transcriptomic data 94%
- Urine single cell RNA-sequencing in focal segmental glomerulosclerosis reveals inflammatory signatures in immune cells and podocytes 93%
- ATP-citrate lyase as a therapeutic target in chronic kidney disease: a Mendelian Randomization analysis 93%
Similar papers in this journal
- Correlating Deep Learning-Based Automated Reference Kidney Histomorphometry with Patient Demographics and Creatinine 93%
- Nephron Number and Kidney Outcomes in IgA Nephropathy: A Retrospective Cohort Study 92%
- Clinical utility of genetic and genomic testing in the precision diagnosis and management of pediatric patients with kidney and urinary tract diseases 91%
Similar papers in this journal
- PERADIGM: Phenotype Embedding Similarity-based Rare Disease Gene Mapping 96%
- Identification Of Modifier Gene Variants Overrepresented In Familial Hypomagnesemia With Hypercalciuria And Nephrocalcinosis Patients With A More Aggressive Renal Phenotype 95%
- Significant Sparse Polygenic Risk Scores across 813 traits in UK Biobank 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.