Back

WikiGOA: Gene set enrichment analysis based on Wikipedia and the Gene Ontology

Lubiana, T.; Dias, T. L.; Peixe, D. G.; Nakaya, H. T. I.

2022-09-17 bioinformatics
10.1101/2022.09.15.508149 bioRxiv
Show abstract

O_LIGene sets curated to Gene Ontology terms are widely used by the transcriptomics community C_LIO_LIPresence in Wikipedia is a common proxy for the relevance of a concept. C_LIO_LIIn this work, we describe the use of Wikidata to generate a dataset comprising only gene sets with a corresponding Wikipedia page. C_LIO_LIWe refer to the dataset as "WikiGOA", standing for "Wikipedia Gene Ontology Annotations" C_LIO_LIWe use the dataset to analyze gene expression data and show that it provides readily understandable results. C_LIO_LIWe envision WikiGOA to be useful for exploring complex biological datasets both in academic research and educational contexts. C_LI NoteThis report was written in a non-standard, experimental format, where assertions are expressed in bullet points. This was done to clarify statements and assumptions, simplify reading and pave the way for conversion to structured formats (e.g., nanopublications). [1]

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.