Real-world data from the Japanese National Health Insurance System enables fine phenotyping in a 14K-scale population-based study
Otsuka-Yamasaki, Y.; Nishiya, N.; Sutoh, Y.; Aoyama, R.; Tanno, K.; Nakao, M.; Komaki, S.; Minabe, S.; Ohmomo, H.; Asahi, K.; Ishigaki, Y.; Sasaki, M.; Shimizu, A.
Show abstract
The use of medical databases, known as real-world data (RWD), enables accurate and efficient estimation of disease prevalence in population-based studies, making it a potential game-changer in epidemiology. However, the lack of standardized data formats across hospitals complicates integration across institutions. In Japan, health insurance claims data are standardized under a nationally unified format, providing a reliable source of structured RWD. We evaluated the utility of insurance claims data in epidemiological research. Incorporating both diagnosis and prescription information into case definitions resulted in four- and six-fold increases in the estimated prevalence of Alzheimers disease (AD) and Parkinsons disease (PD), respectively, compared with conventional self-reported definitions. Subsequent genome-wide association studies (GWAS) for AD showed increased model log-likelihood and identified a characteristic APOE signal, findings observed only with extended case definitions. The APOE effect size was consistent with large case-control studies, while standard errors remained comparable to smaller studies. These results indicate that claims-based phenotyping improves case identification without loss of accuracy and supports scalable approaches for genomic epidemiology and public health surveillance. Patient consent statementAll participants provided written informed consent and the study protocol was approved by the Institutional Review Board of Iwate Medical University (Approval number HG2021-009) Permission to reproduce material from other sourcesNo materials requiring permission from other sources have been used in this manuscript. Clinical trial registrationThis study does not involve interventional components and therefore was not registered as a clinical trial.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting Lung Cancer in Korean Never-Smokers with Polygenic Risk Scores 91%
- Refinement of a published gene-physical activity interaction impacting HDL-cholesterol: role of sex and lipoprotein subfractions 90%
- Genome-wide association studies of 27 accelerometry-derived physical activity measurements identified novel loci and genetic mechanisms 90%
Similar papers in this journal
Similar papers in this journal
- The predictive capacity of polygenic risk scores for disease risk is only moderately influenced by imputation panels tailored to the target population 91%
- Deep-PheWAS: a pipeline for phenotype generation and association analysis for phenome-wide association studies 90%
- TIGA: Target illumination GWAS analytics 89%
Similar papers in this journal
- Biological and Disease Hallmarks of Alzheimer’s Disease Defined by Alzheimer’s Disease Genes 92%
- Trends in cognitive function before and after myocardial infarction : findings from the China Health and Retirement Longitudinal Study 92%
- Plcg2M28L interacts with high fat-high sugar diet to accelerate Alzheimers disease-relevant phenotypes in mice 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.