Back

Leveraging NLP to Identify Domain-Specific Variables in Large-Scale Cohort Metadata: A Sleep Use Case

Draper, B.; Briggs, P.; Purcell, S. M.; Dijk, D.-J.; Bauermeister, S.; Bartsch, U.

2026-01-19 health informatics
10.64898/2026.01.18.26344317 medRxiv
Show abstract

Public health policies increasingly rely on the use of complex and large datasets containing heterogeneous, multimodal data that require advanced analytical methods to extract meaningful insights and support evidence-based decision-making. Essential for the sharing and analysis of public health data is the description of the data ("data about data") or metadata. Indeed, a lack of metadata standards has been identified as a key technical barrier to public health data sharing 1. Metadata varies considerably between cohort studies. It often contains extremely heterogenous variable and data descriptions for similar or identical metrics. The assessment and processing of metadata is very labour intensive. It relies on researchers manually sifting through large numbers of variable descriptors and other study data documentation to extract variables of interest. Unless a validated quantitative tool is used, the variability in phrasing is surprisingly high. Variations in metadata can makes it hard to find and compare results across studies. Specifically, questions about wake-sleep behaviour like subjective sleep quality and duration are commonly employed in large cohort studies - but vary not only in the wording of the questions put to participants but also in the metadata that describes the questions and their responses. We developed a semi-automatic retrieval method (METAMATCH) for sleep-related variables from metadata obtained from multiple large cohort studies. Here we employ sleep as the health area of interest, but in principle this method can be applied to any area of research. We developed the retrieval method using metadata provided by Dementias Platform UK (DPUK, https://www.dementiasplatform.uk/) a Trusted Research Environment (TRE) that hosts over 50 cohort datasets of different sizes. From this we extracted a metadata corpus describing 86,682 variables across 17 cohort studies. To identify sleep-related variables, we curated a reference dictionary (the SLEEPTALKING corpus) combining terms from the Unified Medical Language System (UMLS) and expert knowledge. We then employed Term Frequency-Inverse Document Frequency (TF-IDF) vectorisation and cosine-similarity to rank variable-descriptions by comparing the metadata corpus to the SLEEPTALKING corpus. We identified 337 semantically unique sleep-related variables. We then performed semantic analyses to characterize these variables. We applied Bidirectional Encoder Representations from Transformers (BERT) embeddings and unsupervised cluster analysis. The results indicate that definitions tend to group by cohort and length of the variable description. We validated the cluster analysis using expert-based categorisation of variable descriptions. The most common categories identified were Sleep Difficulty, Sleep Duration, and Sleep Latency & Waking. Consequently, expert knowledge and consensus remains essential for accurately categorizing variables within a universal ontology across different studies. We conclude that combining NLP techniques with expert input offers an efficient approach to harvesting metadata across multiple cohort studies.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.