Leveraging NLP to Identify Domain-Specific Variables in Large-Scale Cohort Metadata: A Sleep Use Case
Draper, B.; Briggs, P.; Purcell, S. M.; Dijk, D.-J.; Bauermeister, S.; Bartsch, U.
Show abstract
Public health policies increasingly rely on the use of complex and large datasets containing heterogeneous, multimodal data that require advanced analytical methods to extract meaningful insights and support evidence-based decision-making. Essential for the sharing and analysis of public health data is the description of the data ("data about data") or metadata. Indeed, a lack of metadata standards has been identified as a key technical barrier to public health data sharing 1. Metadata varies considerably between cohort studies. It often contains extremely heterogenous variable and data descriptions for similar or identical metrics. The assessment and processing of metadata is very labour intensive. It relies on researchers manually sifting through large numbers of variable descriptors and other study data documentation to extract variables of interest. Unless a validated quantitative tool is used, the variability in phrasing is surprisingly high. Variations in metadata can makes it hard to find and compare results across studies. Specifically, questions about wake-sleep behaviour like subjective sleep quality and duration are commonly employed in large cohort studies - but vary not only in the wording of the questions put to participants but also in the metadata that describes the questions and their responses. We developed a semi-automatic retrieval method (METAMATCH) for sleep-related variables from metadata obtained from multiple large cohort studies. Here we employ sleep as the health area of interest, but in principle this method can be applied to any area of research. We developed the retrieval method using metadata provided by Dementias Platform UK (DPUK, https://www.dementiasplatform.uk/) a Trusted Research Environment (TRE) that hosts over 50 cohort datasets of different sizes. From this we extracted a metadata corpus describing 86,682 variables across 17 cohort studies. To identify sleep-related variables, we curated a reference dictionary (the SLEEPTALKING corpus) combining terms from the Unified Medical Language System (UMLS) and expert knowledge. We then employed Term Frequency-Inverse Document Frequency (TF-IDF) vectorisation and cosine-similarity to rank variable-descriptions by comparing the metadata corpus to the SLEEPTALKING corpus. We identified 337 semantically unique sleep-related variables. We then performed semantic analyses to characterize these variables. We applied Bidirectional Encoder Representations from Transformers (BERT) embeddings and unsupervised cluster analysis. The results indicate that definitions tend to group by cohort and length of the variable description. We validated the cluster analysis using expert-based categorisation of variable descriptions. The most common categories identified were Sleep Difficulty, Sleep Duration, and Sleep Latency & Waking. Consequently, expert knowledge and consensus remains essential for accurately categorizing variables within a universal ontology across different studies. We conclude that combining NLP techniques with expert input offers an efficient approach to harvesting metadata across multiple cohort studies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Novel Digital Markers of Sleep Dynamics: A Causal Inference Approach Revealing Age and Gender Phenotypes in Obstructive Sleep Apnea 94%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
- Topographical relocation of adolescent sleep spindles reveals a new maturational pattern of the human brain 92%
Similar papers in this journal
- Sleep in Frontline Healthcare Workers on Social Media During the COVID-19 Pandemic 93%
- Circadian Rhythm Analysis Using Wearable Device Data: A Novel Penalized Machine Learning Approach 92%
- Uncovering social states in healthy and clinical populations using digital phenotyping and Hidden Markov Models 90%
Similar papers in this journal
- The potential of ensemble-based automated sleep staging on single-channel EEG signal from a wearable device 94%
- Looking for a reference for large datasets: relative reliability of visual and automatic sleep scoring 93%
- The effects of daylight saving time clock changes on accelerometer-measured sleep duration in the UK Biobank 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.