Back

Pseudp-p-Value-Based Clumping Enhanced Proteome-wide Mendelian Randomization with Application in Identifying Coronary Heart Disease-Associated Plasma Proteins

Wang, Y.; Cao, Y.; Chen, D.; Shi, D.; Fu, L.; Chen, A.; Shen, S.; Hu, Y.-Q.

2025-01-15 epidemiology
10.1101/2025.01.13.25320450 medRxiv
Show abstract

Mendelian randomization (MR) is a powerful tool for causal inference in epidemiology. However, the presence of weak instrumental variables (IVs) and pleiotropy can lead to biased causal effect estimates. To address these issues, we develop MR-GMM, a novel MR method based on a Gaussian Mixture Model. MR-GMM classifies IVs into four categories--invalid, valid, invalid&null, and null IVs-- and models their effects using a two-dimensional spike-and-slab distribution. Simulation studies demonstrate the high efficiency and robustness of MR-GMM compared to existing methods. More importantly, we propose a pseudo-p-value-based linkage disequilibrium (LD) clumping procedure to address selection bias. This refined procedure is capable of enhancing the performance of MR-GMM as well as many existing MR methods in real-world scenarios. Applying MR-GMM in a large-scale proteome-wide MR study, we identify 45 coronary heart disease-associated plasma proteins. Subsequent network and enrichment analyses highlight the potential of these proteins as biomarkers for disease diagnosis and therapeutic development.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.