Integrative Harmonization of Phenotypic and Genomic Data Improves Bone Mineral Density Prediction in Multi-Study Osteoporosis Research
Liu, A.; Liu, J.; Wu, L.; Wu, Q.
Show abstract
PurposeHarmonizing osteoporosis-related data across multiple datasets is essential for improving the accuracy and generalizability of bone mineral density (BMD) assessments. This study developed a harmonization framework to standardize phenotypic and genomic variables across three major U.S. osteoporosis datasets: GDBF, GWAS, and NHANES. MethodsWe standardized key phenotypic variables (BMD, body mass index (BMI), age, sex, and race/ethnicity) using cohort-specific data dictionaries and applied multiple imputations by chained equations (MICE) to manage missing data. Genomic data were harmonized using principal component analysis (PCA)-based batch effect corrections. Residual regression methods were applied to standardize BMD values. The effectiveness of harmonization on BMD prediction was evaluated using generalized estimating equations (GEE) and mixed-effects models. ResultsPost-harmonization, inter-study variability in BMI was significantly reduced ({Omega}{superscript 2} = 0.0028), and BMD associations with covariates remained consistent across datasets. Harmonized models showed improved predictive performance, with explained variance in BMD increasing (R{superscript 2} = 0.14). PCA confirmed the effective alignment of genetic data, reducing batch effects and improving cross-study compatibility. ConclusionThis study demonstrates the feasibility and effectiveness of harmonizing phenotypic and genomic data for osteoporosis research. The harmonization framework enhances BMD prediction accuracy, supports more inclusive osteoporosis risk assessment, and improves the integration of multi-cohort datasets for future research. These findings highlight the potential of data harmonization in advancing precision medicine for osteoporosis prevention and management. Mini-AbstractHarmonizing osteoporosis datasets improves BMD prediction accuracy and enhances risk assessment. This study developed a harmonization framework integrating phenotypic and genomic data across three major datasets. After harmonization, predictive model performance improved, enabling better osteoporosis risk stratification and advancing precision medicine for fracture prevention.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Leptin signaling and the intervertebral disc: Sex dependent effects of leptin receptor deficiency and Western diet on the spine in a type 2 diabetes mouse model 93%
- Prrx1-driven LINC complex disruption in vivo reduces osteoid deposition but not bone quality after voluntary wheel running 93%
- Moderate Static Magnetic Fields Prevent Estrogen Deficiency-Induced Bone Loss: Evidence from Ovariectomized Mouse Model and Small Sample Size Randomized Controlled Clinical Trial 93%
Similar papers in this journal
- Low Intensity Vibration Protects the Weight Bearing Skeleton and Suppresses Fracture Incidence in Boys with Duchenne Muscular Dystrophy 94%
- Thermoneutral Housing has Limited Effects on Social Isolation-Induced Bone Loss in Male C57BL/6J Mice 93%
- The cortical bone metabolome of C57BL/6J mice is sexually dimorphic 92%
Similar papers in this journal
- Metabolomics insights into osteoporosis through association with bone mineral density 94%
- The gSOS Polygenic Score is Associated with Bone Density and Fracture Risk in Childhood 94%
- The effect of plasma lipids and lipid lowering interventions on bone mineral density: a Mendelian randomization study 94%
Similar papers in this journal
- Sexual dimorphism of MASLD-driven bone loss 95%
- Clinical Observation of Diminished Bone Quality and Quantity through Longitudinal HR-pQCT-derived Remodeling and Mechanoregulation 94%
- Machine Learning Approaches for the Prediction Bone Mineral Density by using genomic and phenotypic data of 5,130 older Men 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.