Back

HostSeq: A Canadian Whole Genome Sequencing and Clinical Data Resource

Yoo, S.; Garg, E.; Elliott, L. T.; Hung, R. J.; Halevy, A. R.; Brooks, J. D.; Bull, S. B.; Gagnon, F.; Greenwood, C. M.; Lawless, J. F.; Paterson, A. D.; Sun, L.; Zawati, M. H.; Lerner-Ellis, J.; Abraham, R. J.; Birol, I.; Bourque, G.; Garant, J.-M.; Gosselin, C.; Li, J.; Whitney, J.; Thiruvahindrapuram, B.; Herbrick, J.-A.; Lorenti, M.; Reuter, M. S.; Liu, S.; Allen, U.; Bernier, F. P.; Biggs, C. M.; Cheung, A. M.; Cowan, J.; Herridge, M.; Maslove, D. M.; Modi, B. P.; Mooser, V.; Morris, S. K.; Ostrowski, M.; Parekh, R. S.; Pfeffer, G.; Suchowersky, O.; Taher, J.; Turvey, S. E.; Upton, J.; Wa

2022-05-10 epidemiology
10.1101/2022.05.06.22274627 medRxiv
Show abstract

HostSeq was launched in April 2020 as a national initiative to integrate whole genome sequencing data from 10,000 Canadians infected with SARS-CoV-2 with clinical information related to their disease experience. The mandate of HostSeq is to support the Canadian and international research communities in their efforts to understand the risk factors for disease and associated health outcomes and support the development of interventions such as vaccines and therapeutics. HostSeq is a collaboration among 13 independent epidemiological studies of SARS-CoV-2 across five provinces in Canada. Aggregated data collected by HostSeq are made available to the public through two data portals: a phenotype portal showing summaries of major variables and their distributions, and a variant search portal enabling queries in a genomic region. Individual-level data is available to the global research community for health research through a Data Access Agreement and Data Access Compliance Office approval. Here we provide an overview of the collective project design along with summary level information for HostSeq. We highlight several statistical considerations for researchers using the HostSeq platform regarding data aggregation, sampling mechanism, covariate adjustment, and X chromosome analysis. In addition to serving as a rich data source, the diversity of study designs, sample sizes, and research objectives among the participating studies provides unique opportunities for the research community.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.