Back

eMIND: Enabling automatic collection of protein variation impacts in Alzheimer's disease from the literature

Gupta, S.; Qin, X.; Wang, Q.; Cowart, J.; Huang, H.; Wu, C. H.; Vijay-Shanker, K.; Arighi, C. N.

2023-09-12 bioinformatics
10.1101/2023.09.07.556602 bioRxiv
Show abstract

Alzheimers disease and related dementias (AD/ADRDs) are among the most common forms of dementia, and yet no effective treatments have been developed. To gain insight into the disease mechanism, capturing the connection of genetic variations to their impacts, at the disease and molecular levels, is essential. The scientific literature continues to be a main source for reporting experimental information about the impact of variants. Thus, development of automatic methods to identify publications and extract the information from the unstructured text would facilitate collecting and organizing information for reuse. We developed eMIND, a deep learning-based text mining system that supports the automatic extraction of annotations of variants and their impacts in AD/ADRDs. In particular, we use this method to capture the impacts of protein-coding variants affecting a selected set of protein properties, such as protein activity/function, structure and post-translational modifications. We conducted an evaluation on the efficacy of eMIND to extract variant impact relations and obtained a recall of 0.84 and a precision of 0.94. The publications and extracted information are integrated into the UniProtKB computationally mapped bibliography to expand annotations on protein entries. eMINDs text-mined output are presented using controlled vocabularies and ontologies for variant, disease and impact along with the evidence sentences. A sample of annotated abstracts can be accessed at URL: https://research.bioinformatics.udel.edu/itextmine/emind.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.