WikiMedMap: Expanding the Phenotyping Mapping Toolbox Using Wikipedia
Sulieman, L.; Wu, P.; Denny, J.; Bastarache, L.
Show abstract
Researchers utilizing phenotypic data from diverse sources require matching of phenotypes to standard clinical vocabularies. Mapping phenotypes to vocabulary can be difficult, as existing tools are often incomplete, can be difficult to access, and can be cumbersome to use, especially for non-experts. We created WikiMedMap as a simple tool that leverages Wikipedia and maps phenotype strings to standard clinical vocabularies. We assessed WikiMedMap by mapping phenotype strings from questionnaires in the UK Biobank and from Mendelian diseases in Online Mendelian Inheritance in Man (OMIM) database to eight vocabularies: International Classification of Diseases, Ninth Revision (ICD-9), ICD-10, ICD-O, Medical Subject Headings (MeSH), OMIM, Disease Database, and MedlinePlus. WikiMedMap outperformed conventional mapping tools in finding potential matches for phenotype strings. We envision WikiMedMap as a technique that complements existing and established tools to map strings to clinical vocabularies that usually do not coexist in one source.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- PheMIME: An Interactive Web App and Knowledge Base for Phenome-Wide, Multi-Institutional Multimorbidity Analysis 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 94%
Similar papers in this journal
Similar papers in this journal
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 94%
- Development of a COVID-19 Application Ontology for the ACT Network 94%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.