HostSeq: A Canadian Whole Genome Sequencing and Clinical Data Resource
Yoo, S.; Garg, E.; Elliott, L. T.; Hung, R. J.; Halevy, A. R.; Brooks, J. D.; Bull, S. B.; Gagnon, F.; Greenwood, C. M.; Lawless, J. F.; Paterson, A. D.; Sun, L.; Zawati, M. H.; Lerner-Ellis, J.; Abraham, R. J.; Birol, I.; Bourque, G.; Garant, J.-M.; Gosselin, C.; Li, J.; Whitney, J.; Thiruvahindrapuram, B.; Herbrick, J.-A.; Lorenti, M.; Reuter, M. S.; Liu, S.; Allen, U.; Bernier, F. P.; Biggs, C. M.; Cheung, A. M.; Cowan, J.; Herridge, M.; Maslove, D. M.; Modi, B. P.; Mooser, V.; Morris, S. K.; Ostrowski, M.; Parekh, R. S.; Pfeffer, G.; Suchowersky, O.; Taher, J.; Turvey, S. E.; Upton, J.; Wa
Show abstract
HostSeq was launched in April 2020 as a national initiative to integrate whole genome sequencing data from 10,000 Canadians infected with SARS-CoV-2 with clinical information related to their disease experience. The mandate of HostSeq is to support the Canadian and international research communities in their efforts to understand the risk factors for disease and associated health outcomes and support the development of interventions such as vaccines and therapeutics. HostSeq is a collaboration among 13 independent epidemiological studies of SARS-CoV-2 across five provinces in Canada. Aggregated data collected by HostSeq are made available to the public through two data portals: a phenotype portal showing summaries of major variables and their distributions, and a variant search portal enabling queries in a genomic region. Individual-level data is available to the global research community for health research through a Data Access Agreement and Data Access Compliance Office approval. Here we provide an overview of the collective project design along with summary level information for HostSeq. We highlight several statistical considerations for researchers using the HostSeq platform regarding data aggregation, sampling mechanism, covariate adjustment, and X chromosome analysis. In addition to serving as a rich data source, the diversity of study designs, sample sizes, and research objectives among the participating studies provides unique opportunities for the research community.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 92%
- Researching COVID to enhance recovery (RECOVER) pediatric study protocol: Rationale, objectives and design 92%
- SARS-CoV-2 Omicron subvariant genomic variation associations with immune evasion in Northern California: A retrospective cohort study 92%
Similar papers in this journal
- Cohort Profile: Post-hospitalisation COVID-19 study (PHOSP-COVID) 92%
- An empirical investigation into the impact of winner's curse on estimates from Mendelian randomization 90%
- Cohort profile: Virus Watch: Understanding community incidence, symptom profiles, and transmission of COVID-19 in relation to population movement and behaviour 90%
Similar papers in this journal
- HeartBioPortal2.0: new developments and updates for genetic ancestry and cardiometabolic quantitative traits in diverse human populations 92%
- Color Data v2: a user-friendly, open-access database with hereditary cancer and hereditary cardiovascular conditions datasets 91%
- HumanMine: advanced data searching, analysis and cross-species comparison. 90%
Similar papers in this journal
- Collaborative Cohort of Cohorts for COVID-19 Research (C4R) Study: Study Design 95%
- A system for phenotype harmonization in the NHLBI Trans-Omics for Precision Medicine (TOPMed) Program 93%
- Potential Biases in Test-Negative Design Studies of COVID-19 Vaccine Effectiveness Arising from the Inclusion of Asymptomatic Individuals 91%
Similar papers in this journal
- Use of >100,000 NHLBI Trans-Omics for Precision Medicine (TOPMed) Consortium whole genome sequences improves imputation quality and detection of rare variant associations in admixed African and Hispanic/Latino populations 93%
- Proteome-wide Mendelian randomization identifies causal links between blood proteins and severe COVID-19 93%
- Evaluation of Bayesian Linear Regression Models for Gene Set Prioritization in Complex Diseases 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.