Back

Knowledge representation of a multi-centre adolescent and young adult (AYA) cancer infrastructure; development of the STRONG AYA Knowledge Graph

Hogenboom, J.; Gouthamchand, V.; Cairns, C.; Janssen, S. H. M.; Way, K.; Dekker, A. L. A. J.; van der Graaf, W. T. A.; Darlington, A.-S.; Husson, O.; Wee, L. Y. L.; van Soest, J.; Lobo Gomes, A.

2025-06-03 health informatics
10.1101/2025.06.03.25328788 medRxiv
Show abstract

PurposeRare diseases are difficult to fully capture, and regularly call for large, geographically dispersed initiatives. Such initiatives are often met with data harmonisation challenges. These challenges render data incompatible and impede successful realisation. The STRONG AYA project is such an initiative, specifically focusing on adolescents and young adults (AYAs) with cancer. STRONG AYA is setting up a federated data infrastructure containing data of varying format. Here, we elaborate on how we used healthcare-agnostic Semantic Web technologies to overcome such challenges. MethodologyWe structured the STRONG AYA case-mix and core outcome measures concepts and their properties as knowledge graphs. Having identified the corresponding standard terminologies, we developed a semantic map based on the knowledge graphs and the here introduced annotation helper plugin for Flyover. Flyover is a tool that converts structured data into Resource Descriptor Framework (RDF) triples and enables semantic interoperability. As a demonstration, we mapped data that is to be included in the STRONG AYA infrastructure. ResultsThe knowledge graphs provided a comprehensive overview of the large number of STRONG AYA concepts. The semantic terminology mapping and annotation helper allowed us to query data with incomprehensible terminologies, without changing them. Both the knowledge graphs and semantic map were made available on a Hugo webpage for increased transparency and understanding. DiscussionThe use of Semantic Web technologies such as RDF and knowledge graphs are a viable solution to overcome challenges regarding data interoperability and reusability for a federated AYA cancer data infrastructure without being bound to rigid standardised schemas. The linkage of semantically meaningful concepts to otherwise incomprehensible data elements demonstrates how by using these domain-agnostic technologies we made non-standardised healthcare data interoperable.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.