GWAS Summary Statistic Tool: A Meta-Analysis and Parsing Tool for Polygenic Risk Score Calculation
Muneeb, M. -; Ascher, D.
Show abstract
MotivationGWAS (genome-wide association study) summary statistic files are essential inputs for polygenic risk score (PRS) calculation, yet identifying suitable files across thousands of catalog entries requires downloading large files and manually inspecting their column structures--a process that is time-consuming and storage-intensive. ResultsWe present GWASPoker, a phenotype-driven, GWAS-Catalog-specific pre-download triage tool that scans candidate GWAS files for PRS column availability by partial download and header detection, without requiring full-file transfer. Analysing 60,499 records from the GWAS Catalog, 60,281 (99.6%) contained accessible download links, of which 54,026 (89.6%) were successfully partially downloaded and parsed across 20 file formats, yielding 724 unique header signatures. Across 13 phenotypes, 84 of 85 manually curated GWAS files (98.8%) were automatically retrieved and processed. Header validation against fully downloaded files showed exact agreement in 23 of 28 cases (82.1%). Availability and implementationGWASPoker is implemented in Python 3 and freely available at https://github.com/MuhammadMuneeb007/GWASPokerforPRS under the MIT licence. Example outputs and documentation are provided in the repository. The tool was tested on Linux (HPC cluster) with Python 3.8+. The LLM-based code-generation step is entirely optional; a rules-based column-mapping template is provided for fully offline use.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DISEASES 2.0: a weekly updated database of disease-gene associations from text mining and data integration 93%
- Curation of over 10,000 transcriptomic studies to enable data reuse 93%
- HeartBioPortal2.0: new developments and updates for genetic ancestry and cardiometabolic quantitative traits in diverse human populations 93%
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 94%
- SnpHub: an easy-to-set-up web server framework for exploring large-scale genomic variation data in the post-genomic era with applications in wheat 94%
- DivBrowse - interactive visualization and exploratory data analysis of variant call matrices 94%
Similar papers in this journal
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 93%
- Bioinformatics workflows for genomic analysis of tumors from Patient Derived Xenografts (PDX): challenges and guidelines 92%
- Accuracy and Reproducibility of Somatic Point Mutation Calling in Clinical-Type Targeted Sequencing Data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.