Back

Exploring biomedical records through text mining-driven complex data visualisation

Pita Costa, J.; Stopar, L.; Rei, L.; Massri, B.; Grobelnik, M.

2021-03-29 health informatics
10.1101/2021.03.27.21250248 medRxiv
Show abstract

The recent events in health call for the prioritization of insightful and meaningful information retrieval from the fastly growing pool of biomedical knowledge. This information has its own challenges both in the data itself and in its appropriate representation, enhancing its usability by health professionals. In this paper we present a framework leveraging the MEDLINE dataset and its controlled vocabulary, the MeSH Headings, to annotate and explore health-related documents. The MEDijs system ingests and automatically annotates text documents, extending their legacy metadata with MeSH Headings. It then uses text mining algorithms that enable interactive data visualisations. These allow the user to the exploration of the enriched data made available by the MEDijs system. CCS CONCEPTS* Information systems; * Computing methodologies [->] Machine learning approaches; ACM Reference FormatJoao Pita Costa, Luka Stopar, Luis Rei, Besher Massri, and Marko Grobelnik. 2018. Exploring biomedical records through text mining-driven complex data visualisation. In Proceedings of SEBILAN 21: ACM International Workshop on Semantics-enabled Biomedical Literature Analytics (SEBILAN 21). ACM, New York, NY, USA, 6 pages. https://doi.org/0

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.