Beyond the classical plasma secretome: genetic architecture and disease associations of the expanded human plasma proteome in 13,445 Europeans
Cadiou, S.; Konig, E.; Mapelli, A.; Pontali, G.; Ghasemi-Semeskandeh, D.; Filosi, M.; Ferolito, B. R.; Massi, M. C.; Cuccuru, G.; Jiang, X.; Winicki, G.; Navarro-Gallinad, A.; Gravel-Pucillo, K.; Pirastu, N.; Landini, A.; Sharapov, S.; Rainer, J.; Gogele, M.; Lundin, R.; Mascalzoni, D.; Biasiotto, R.; Blankenburg, H.; De Grandi, A.; Egger, C.; Arend, L.; Woller, F.; Ieva, F.; Soranzo, N.; Cho, K.; Gaziano, J. M.; Zuccolo, L.; Domingues, F. S.; Pattaro, C.; Danesh, J.; Pramstaller, P. P.; Pereira, A.; Di Angelantonio, E.; Fuchsberger, C.; Giambartolomei, C.; Butterworth, A. S.
Show abstract
Circulating plasma proteins are key biomarkers and therapeutic targets, now measurable at scale through high-throughput technologies, yet whether expanding proteomics platforms beyond the classical plasma secretome enhances genetic discovery and causal inference remains poorly understood. Here, we use an expanded SomaScan 7k platform to map the genetic architecture of a broader segment of the plasma proteome and to evaluate how proteome expansion affects pQTL discovery, causal inference and therapeutic target prioritisation. After quality control, we analysed 7,144 aptamers targeting 6,267 proteins in the harmonised dataset of two European cohorts: INTERVAL (n = 9,251 participants) and CHRIS (n = 4,194), and conducted genome-wide pQTL association analyses followed by meta-analysis. We identified 7,870 significant pQTLs (P-value < 1.26 x 10E-11; 1,784 cis, 6,086 trans), of which 2,704 (34%) associations were not reported in five prior large-scale pQTL studies. Newly assessed proteins, which accounted for 53% (1,422/2,704) of the novel associations, were less likely to harbour cis-pQTLs associations (15%) than those in the previous platform version (28%), consistent with their lower expected plasma concentrations and predominantly intracellular localisation. Colocalization analyses revealed widespread sharing of genetic signals across proteins and characterised 22 pleiotropic trans-regulatory hotspots accounting for 68% of all trans-pQTLs. Through two-sample Mendelian randomization analyses on 2,003 phenotypes from the Million Veteran Program, UK Biobank, and FinnGen (combined N > 1.2 million), we identified 6,340 genetically supported protein-trait associations, highlighting disease mechanisms and potential therapeutic opportunities beyond currently drug-targeted circulating proteins. Together, these findings provide a systematic view of the genetic architecture of the expanded plasma proteome and demonstrate that plasma proteome expansion reveals genetically anchored disease biology beyond the classical secretome, while exposing inherent biological and technical constraints of studying low-abundance intracellular proteins in circulation.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Whole genome sequencing analysis of the cardiometabolic proteome 96%
- A genome-wide association study in 10,000 individuals links plasma N-glycome to liver disease and anti-inflammatory proteins 96%
- Systematic discovery of gene-environment interactions underlying the human plasma proteome in UK Biobank 96%
Similar papers in this journal
- Proteome-wide Mendelian randomization in global biobank meta-analysis reveals multi-ancestry drug targets for common diseases 96%
- The functional impact of rare variation across the regulatory cascade 96%
- Impact of disease-associated chromatin accessibility QTLs across immune cell types and contexts 95%
Similar papers in this journal
- Interaction molecular QTL mapping discovers cellular and environmental modifiers of genetic regulatory effects 95%
- A multi-omic integrative scheme characterizes tissues of action at loci associated with type 2 diabetes 95%
- Characterization of non-coding variants associated with transcription factor binding through ATAC-seq-defined footprint QTLs in liver 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.