Disease prevalence, health-related and socio-demographic factors in the GCAT cohort. A comparison with the general population of Catalonia.
Blay, N.; Carrasco-Ribelles, L. A.; Farre, X.; Iraola-Guzman, S.; Danes-Castells, M.; Violan, C.; de Cid, R.
Show abstract
BackgroundPopulation-based cohorts play a key role in epidemiological studies. However, it is known that volunteer cohorts include a healthy volunteer bias. Assessment and characterization of this bias is needed to extrapolate results to the general population. Here, we assess the bias of the population-based cohort GCAT, encompassing 20 000 adult participants from Catalonia with electronic health record data. The aim of this study is to compare the GCAT cohort with its age-matching Catalan population, to assess their representativeness, as well as determining the weights to make results generalisable. MethodsStatistical comparisons until 2019 in multiple variables across sociodemographic, lifestyle, diseases and medication domains were performed by stratified analysis with Fishers exact test and t-test. Electronic health records of Catalonia (SIDIAP), and registers from the statistics institute of Catalonia (IDESCAT) and Spain (INE) were used to make the comparisons. We generated weights accounting for sociodemographic, lifestyle and multimorbidity factors. ResultsGCAT cohort is enriched in women and younger individuals, with higher socioeconomic status, more health conscious and healthier in terms of mortality and chronic disease prevalence. We have shown that this bias can be corrected with weighting techniques, providing a more representative sample of the general population. ConclusionsThe application of multidomain weights, encompassing not only sociodemographic aspects, but also lifestyle and health-related variables, has effectively diminished the observed bias in disease prevalence estimates within the GCAT cohort. This correction has led to an enhancement of the cohorts representativeness, rendering it more akin to the general population of Catalonia.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and validation of a clinical risk score to predict the risk of SARS-CoV-2 infection from administrative data: a population-based cohort study from Italy 94%
- Characteristics of a tattooed population and a possible role of tattoos as a risk factor for chronic diseases: Results from the LIFE-Adult-Study 94%
- Pandemic trends in health care use: From the hospital bed to the general practitioner with COVID-19 93%
Similar papers in this journal
- COVID19-related and all-cause mortality among middle-aged and older adults across the first epidemic wave of SARS-COV-2 infection in the region of Tarragona, Spain: results from the COVID19 TARRACO Cohort Study, March-June 2020 94%
- Awareness, knowledge and trust in the Greek authorities towards COVID-19 pandemic: results from the Epirus Health Study cohort 94%
- Cognitive impairment at older ages among 8000 men and women living in Mexico City: cross-sectional analyses of a prospective study 92%
Similar papers in this journal
- Cancer and the risk of COVID-19 diagnosis, hospitalisation, and death: a population-based multi-state cohort study including 4,618,377 adults in Catalonia, Spain 93%
- Cumulative COVID-19 incidence, mortality, and prognosis in cancer survivors: a population-based study in Reggio Emilia, Northern Italy 93%
- Associations of plasma omega-6 and omega-3 fatty acids with overall and 19 site-specific cancers: a population-based cohort study in UK Biobank 92%
Similar papers in this journal
- Real world evidence of calcifediol use and mortality rate of COVID-19 hospitalized in a large cohort of 16,401 Andalusian patients 94%
- Development and performance of a population-based risk stratification model for COVID-19 94%
- Understanding multimorbidity trajectories in Scotland: an application of sequence analysis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.