Back

Development and evaluation of codelists for identifying marginalised groups in primary care

Perchyk, T.; de Vere Hunt, I.; Nicholson, B. D.; Mounce, L.; Sykes, K.; Lyratzopoulos, Y.; Lemanska, A.; Whitaker, K. L.; Kerrison, R. S.

2024-09-13 primary care research
10.1101/2024.09.11.24313391 medRxiv
Show abstract

BackgroundPrimary care electronic health records provide a rich source of information for inequalities research. However, the reliability and validity of the research derived from these records depends on the completeness and resolution of the codelists used to identify marginalised populations. AimThe aim of this project was to develop comprehensive codelists for identifying ethnic minorities, people with learning disabilities (LD), people with severe mental illness (SMI) and people who are transgender. Design and settingThis study was a codelist development project, conducted using primary care data from the United Kingdom. MethodGroups of interest were defined a priori. Relevant clinical codes were identified by searching Clinical Practice Research Datalink (CPRD) publications, codelist repositories and the CPRD code browser. Relevant codelists were downloaded and merged according to marginalised group. Duplicates were removed and remaining codes reviewed by two general practitioners. Comprehensiveness was assessed in a representative CPRD population of 10,966,759 people, by comparing the frequencies of individuals identified when using the curated codelists, compared to commonly used alternatives. ResultsA total of 52 codelists were identified. 1,420 unique codes were selected after removal of duplicates and GP review. Compared with comparator codelists, an additional 48,017 (76.6%), 52,953 (68.9%) and 508 (36.9%) people with a LD, SMI or transgender code were identified. The frequencies identified for ethnicity were consistent with expectations for the UK population. ConclusionThe codelists curated through this project will improve inequalities research by improving standards of identifying marginalised groups in primary care data. HOW THIS FITS INO_LIThe reliability and validity of primary care data for inequalities research depends on the comprehensiveness of the codes used to identify people from marginalised groups. C_LIO_LIThis study set out to develop comprehensive codelists for the identification of four key groups, known to experience health inequalities. C_LIO_LIWe developed comprehensive codelists for identifying ethnic minorities, learning disabilities, severe mental illness and people who are transgender, using a systematic approach. C_LIO_LIThe codelists were validated by two general practitioners, assessed in a representative sample, and can now be used in primary care practice and research, both nationally and internationally. C_LI

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.