Illuminating the Druggable Genome through Patent Bioactivity Data
Leach, A. R.; Magarinos, M. P.; Gaulton, A.; Felix, E.; Kiziloren, T.; Arcila, R.; Oprea, T. I.
Show abstract
The patent literature is a potentially valuable source of bioactivity data. The SureChEMBL database (https://www.surechembl.org/) is a publicly available large-scale resource that contains compounds extracted on a daily basis from the full text, images and attachments of patent documents, through an automated text and image-mining pipeline. In this paper we describe a process to prioritise 3.7 million life science relevant patents obtained from SureChEMBL, according to how likely they were to contain bioactivity data for potent small molecules on less-studied targets, according to the classification developed by the Illuminating the Druggable Genome (IDG) project. The overall goal was to select a smaller number of patents that could be manually curated and incorporated into the ChEMBL database. We describe the approach taken, the results obtained, and provide some illustrative examples.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Drug repositoning or target repositioning: a structural perspective of drug-target-indication relationship for available repurposed drugs. 96%
- HerbComb: an integrated database for the discovery of novel combinational therapies from herbal medicines 94%
- Intense bitterness of molecules: machine learning for expediting drug discovery 93%
Similar papers in this journal
- Minimal information for Chemosensitivity assays (MICHA): A next-generation pipeline to enable the FAIRification of drug screening experiments 96%
- Data-driven strategies for drug repurposing: insights, recommendations, and case studies 96%
- Revolutionizing GPCR-Ligand Predictions: DeepGPCR with experimental Validation for High-Precision Drug Discovery 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.