Accurate Name Entity Recognition for Biomedical Literatures: A Combined High-quality Manual Annotation and Deep-learning Natural Language Processing Study
Huang, D.-L.; Zeng, Q.; Xiong, Y.; Liu, S.; Pang, C.; Xia, M.; Fang, T.; Ma, Y.; Qiang, C.; Zhang, Y.; Zhang, Y.; Li, H.; Yuan, Y.
Show abstract
A combined high-quality manual annotation and deep-learning natural language processing study is reported to make accurate name entity recognition (NER) for biomedical literatures. A home-made version of entity annotation guidelines on biomedical literatures was constructed. Our manual annotations have an overall over 92% consistency for all the four entity types -- gene, variant, disease and species --with the same publicly available annotated corpora from other experts previously. A total of 400 full biomedical articles from PubMed are annotated based on our home-made entity annotation guidelines. Both a BERT-based large model and a DistilBERT-based simplified model were constructed, trained and optimized for offline and online inference, respectively. The F1-scores of NER of gene, variant, disease and species for the BERT-based model are 97.28%, 93.52%, 92.54% and 95.76%, respectively, while those for the DistilBERT-based model are 95.14%, 86.26%, 91.37% and 89.92%, respectively. The F1 scores of the DistilBERT-based NER model retains 97.8%, 92.2%, 98.7% and 93.9% of those of BERT-based NER for gene, variant, disease and species, respectively. Moreover, the performance for both our BERT-based NER model and DistilBERT-based NER model outperforms that of the state-of-art model--BioBERT, indicating the significance to train an NER model on biomedical-domain literatures jointly with high-quality annotated datasets.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TCMM: A Unified Database for Traditional Chinese Medicine Modernization and Therapeutic Innovations 94%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 93%
- iMDA-BN: Identification of miRNA-Disease Associations based on the Biological Network and Graph Embedding Algorithm 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 92%
- Alzheimer Disease Knowledge Graph Enhances Knowledge Discovery and Disease Prediction 92%
- Decoding protein binding landscape on circular RNAs with base-resolution Transformer models 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.