Back

Monogenetic Rare Diseases in Biomedical Databases and Text Mining

Nesterova, A. P.; Klimov, E.; Sozin, S.; Sobolev, V.; Linsley, P.; Golovatenko-Abramov, P. K.

2022-04-16 medical education
10.1101/2022.04.07.22273575 medRxiv
Show abstract

1AO_SCPLOWBSTRACTC_SCPLOWThe testing of pharmacological hypotheses becomes faster and more accurate, but at the same time more difficult than even two decades ago. It takes more time to collect and analyse disease mechanisms and experimental facts in various specialized resources. We discuss a new approach to aggregating individual pieces of information about a single disease using Elseviers automated text mining technology. Developed algorithm allows for the collection of published facts in a unified format starting only with the name of the disease. The special template, which combines research and clinical descriptions of diseases was developed. The approach was tested, and information was collected for 55 rare monogenic diseases. Clinical, molecular, and pharmacological characteristics of diseases with supporting references from the literature are available in the form of tables and files. Manually curated templates for 10 rare diseases, including top ranked Cystic Fibrosis and Huntingtons disease, were published to demonstrate the results of the described approach.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.