Back

Study Data Element Mapping: Feasibility of Defining Common Data Elements Across COVID-19 Studies

Mathewson, P.; Gordon, B.; Snowley, K.; Fennessy, C.; Denniston, A.; Sebire, N.

2020-05-26 health informatics
10.1101/2020.05.19.20106641 medRxiv
Show abstract

BackgroundNumerous clinical studies are now underway investigating aspects of COVID-19. The aim of this study was to identify a selection of national and/or multicentre clinical COVID-19 studies in the United Kingdom to examine the feasibility and outcomes of documenting the most frequent data elements common across studies to rapidly inform future study design and demonstrate proof-of-concept for further subject-specific study data element mapping to improve research data management. Methods25 COVID-19 studies were included. For each, information regarding the specific data elements being collected was recorded. Data elements collated were arbitrarily divided into categories for ease of visualisation. Elements which were most frequently and consistently recorded across studies are presented in relation to their relative commonality. ResultsAcross the 25 studies, 261 data elements were recorded in total. The most frequently recorded 100 data elements were identified across all studies and are presented with relative frequencies. Categories with the largest numbers of common elements included demographics, admission criteria, medical history and investigations. Mortality and need for specific respiratory support were the most common outcome measures, but with specific studies including a range of other outcome measures. ConclusionThe findings of this study have demonstrated that it is feasible to collate specific data elements recorded across a range of studies investigating a specific clinical condition in order to identify those elements which are most common among studies. These data may be of value for those establishing new studies and to allow researchers to rapidly identify studies collecting data of potential use hence minimising duplication and increasing data re-use and interoperability

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.