Bringing PanglaoDB to 5-star Linked Open Data using Wikidata
Lubiana, T.; Cavalcante, J. V. F.
Show abstract
PanglaoDB is a database of cell-type markers widely used for single-cell RNA sequencing data analysis. However, cell types and genes in the database are encoded by free text, lacking proper identifiers. Wikidata, is a freely editable knowledge graph database useful for integrating biomedical knowledge. We thus reasoned that porting PanglaoDBs markers to the platform could improve their reusability and overall technical quality (FAIRness). We mapped 188 cell types from PanglaoDB to species-neutral terms on Wikidata and created 376 species-specific terms for cell types in Homo sapiens and Mus musculus. These terms were enriched with marker information via the has marker (P8872) property, totaling over 15.000 cell type X marker associations (w.wiki/9iw6). We explored this new subset of the graph via SPARQL queries, illustrating the discovery potential of structured, integrated knowledge. For example, we found a previously unexplored link between rosehip neurons, clozapine, and schizophrenia via the HRH1 marker. Besides the graph-based insights, we took time to describe the details of the reconciliation process, hoping to stimulate more resources for a move to a 5-star linked open data format.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- CoNECo: A Corpus for Named Entity recognition and normalization of protein Complexes 94%
- MSABrowser: dynamic and fast visualization of sequence alignments, variations, and annotations 93%
- Understanding Ecological Systems Using Knowledge Graphs: An Application to Highly Pathogenic Avian Influenza 93%
Similar papers in this journal
- dialogi: Utilising NLP with chemical and disease similarities to drive the identification of Drug-Induced Liver Injury literature 93%
- Impute.me: an open source, non-profit tool for using data from DTC genetic testing to calculate and interpret polygenic risk scores. 92%
- SCSA: a cell type annotation tool for single-cell RNA-seq data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.