Named Entity Recognition of Pharmacokinetic parameters in the scientific literature
Hernandez, F. G.; Nguyen, Q.; Smith, V. C.; Cordero, J. A.; Ballester, M. R.; Duran, M.; Sole, A.; Chotsiri, P.; Wattanakul, T.; Mundin, G.; Lilaonitkul, W.; Standing, J. F.; Kloprogge, F.
Show abstract
The development of accurate predictions for a new drugs absorption, distribution, metabolism, and excretion profiles in the early stages of drug development is crucial due to high candidate failure rates. The absence of comprehensive, standardised, and updated pharmacokinetic (PK) repositories limits pre-clinical predictions and often requires searching through the scientific literature for PK parameter estimates from similar compounds. While text mining offers promising advancements in automatic PK parameter extraction, accurate Named Entity Recognition (NER) of PK terms remains a bottleneck due to limited resources. This work addresses this gap by introducing novel corpora and language models specifically designed for effective NER of PK parameters. Leveraging active learning approaches, we developed an annotated corpus containing over 4,000 entity mentions found across the PK literature on PubMed. To identify the most effective model for PK NER, we fine-tuned and evaluated different NER architectures on our corpus. Fine-tuning BioBERT exhibited the best results, achieving a strict F1 score of 90.37% in recognising PK parameter mentions, significantly outperforming heuristic approaches and models trained on existing corpora. To accelerate the development of end-to-end PK information extraction pipelines and improve pre-clinical PK predictions, the PK NER models and the labelled corpus were released open source at https://github.com/PKPDAI/PKNER.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Automated recognition of functional compound-protein relationships in literature 95%
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 92%
- A cautionary tale about properly vetting datasets used in supervised learning predicting metabolic pathway involvement 92%
Similar papers in this journal
- A fast, accurate, and generalisable heuristic-based negation detection algorithm for clinical text 94%
- Alzheimer Disease Knowledge Graph Enhances Knowledge Discovery and Disease Prediction 90%
- From Web to RheumaLpack: Creating a Linguistic Corpus for Exploitation and Knowledge Discovery in Rheumatology 90%
Similar papers in this journal
- Multiscale virtual screening optimization for shotgun drug repurposing using the CANDO platform 89%
- VAE-Sim: a novel molecular similarity measure based on a variational autoencoder 89%
- Identifying Protein Features and Pathways Responsible for Toxicity using Machine learning, CANDO, and Tox21 datasets: Implications for Predictive Toxicology 89%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.