Back

UKBCC: a cohort curation package for UK Biobank

Kiral, I.; Willems, N.; Goudey, B.

2020-07-13 bioinformatics
10.1101/2020.07.12.199810 bioRxiv
Show abstract

SummaryThe UK Biobank (UKB) has quickly become a critical resource for researchers conducting a wide-range of biomedical studies (Bycroft et al., 2018). The database is constructed from heterogeneous data sources, employs several different encoding schemes, and is disparately distributed throughout UKB servers. Consequently, querying these data remains complicated, making it difficult to quickly identify participants who meet a given set of criteria. We have developed UK Biobank Cohort Curator (UKBCC), a Python tool that allows researchers to rapidly construct cohorts based on a set of search terms. Here, we describe the UKBCC implementation, critical sub-modules and functions, and outline its usage through an example use case for replicable cohort creation. AvailabilityUKBCC is available through PyPi (https://pypi.org/project/ukbcc) and as open source code on GitHub (https://github.com/tool-bin/ukbcc). Contactisa.kiral@gmail.com

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.