From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores
Barbosa Araujo, P. V.; da Silva Fiuza, T.; Ferraz, R. S.; Kroll, J. E.; Andrade, R. L.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; de Souza, S. J.
Show abstract
Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- From Text to Translation: Using Language Models to Prioritize Variants for Clinical Review 94%
- Comparative analysis of functional assay evidence use by ClinGen Variant Curation Expert Panels 93%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 94%
- Identification of allele-specific KIV-2 repeats and impact on Lp(a) measurements for cardiovascular disease risk 92%
- Identification of single nucleotide variants using position-specific error estimation in deep sequencing data 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.