Back

Whole genome sequence meta-analyses reveal common and rare genetic associations with critical COVID-19

Kousathanas, A.; Rawlik, K.; Pairo-Castineira, E.; Griffiths, F.; Oosthuyzen, W.; Clohisey Hendry, S.; Malinauskas, T.; Butler-Laporte, G.; Arumugam, P.; Begg, C.; Chadeau-Hyam, M.; Chan, G.; Cooke, G.; Donovan, S.; Elgar, G.; Fowler, T. A.; Goddard, P.; Hinds, C.; Horby, P.; Ling, L.; Magavern, E. F.; Maleady-Crowe, F.; Montgomery, H.; Odhams, C. A.; Openshaw, P. J. M.; Patch, C.; Rendon, A.; Salehi, S.; Scott, R. H.; Semple, M. G.; Shankar-Hari, M.; Siddiq, A.; Stuckey, A.; Summers, C.; Todd, L.; Walker, S.; Walsh, T.; Ward, H.; Zainy, T.; GenOMICC Investigators, ; ISARIC4C Investigators,

2025-11-21 intensive care and critical care medicine
10.1101/2025.11.19.25340573 medRxiv
Show abstract

In susceptible patients, COVID-19 causes life-threatening disease driven by immune-mediated inflammatory lung injury. We have previously shown that multiple common host genetic variants are significantly associated with susceptibility to critical Covid-19, 1;2;3 and in one case, we demonstrated that such variants can inform development of new, effective drug treatment1;4. Here we report an association analysis of whole-genome sequences (WGS) from 11,423 cases from the GenOMICC study and 60,628 controls, together with meta-analyses with available genome-wide data (Fig. 1). We identify a rare association signal at SLC50A1, primarily driven by a missense variant rs147850817 (1:155138217:G:T, Arg201Leu) that may interfere with transport function, and we identify four common association signals near ARF1, ZNF462, KLF13 and MVP genes. Finally, we build a WGS-derived polygenic risk score (PRS) for critical Covid-19, which offers only marginal improvement in risk estimation for the general population but may provide clinically-valuable discrimination for extreme susceptibility. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=137 SRC="FIGDIR/small/25340573v1_fig1.gif" ALT="Figure 1"> View larger version (37K): org.highwire.dtl.DTLVardef@6c4561org.highwire.dtl.DTLVardef@3efb16org.highwire.dtl.DTLVardef@d68561org.highwire.dtl.DTLVardef@1ceccbd_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig. 1:C_FLOATNO Overview of analysed cohorts and study analysis pipeline. The analysed cohorts comprised of individuals with severe COVID-19 and individuals with mild COVID-19 symptoms, along with individuals from the 100,000 Genomes Project (100kGP). COVID-19 severe and mild cohorts and a subset of 100kGP individuals were processed using the Genomics England Pipeline 2.0 (Illumina Dragen) with an additional subset of 100kGP individuals processed using a different pipeline (Illumina NSV4). Two separate aggregates were merged after masking of low quality genotypes, followed by sample quality control (sample-QC) for relatedness, sex mismatches and sample level quality. The post-quality control (post-QC) samples were divided into COVID-19 severe cases and three distinct control groups: ctrl-mld, ctrl-dgn, and ctrl-all and the sample breakdown by ancestry is shown. The ctrl-all controls set was used for GWAS analyses, while the ctrl-dgn controls set was used for rare variant aggregate testing (RVAT) analyses, as potential noise from inclusion of data from different processing pipelines is more challenging to quality control for rare variants. The ctrl-mld control set was utilized in sensitivity analyses to validate the primary findings. Additional site-wise quality control appropriate for GWAS and RVAT analyses was performed, followed by meta-analysis with other studies. A breakdown of case and control samples across studies that were included in the GWAS and RVAT meta-analyses is shown. C_FIG

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.