A Comprehensive Benchmark Study on Biomedical Text Generation and Mining with ChatGPT
Chen, Q.; Sun, H.; Liu, H.; Jiang, Y.; Ran, T.; Jin, X.; Xiao, X.; Lin, Z.; Niu, Z.; Chen, H.
Show abstract
In recent years, the development of natural language process (NLP) technologies and deep learning hardware has led to significant improvement in large language models(LLMs). The ChatGPT, the state-of-the-art LLM built on GPT-3.5, shows excellent capabilities in general language understanding and reasoning. Researchers also tested the GPTs on a variety of NLP related tasks and benchmarks and got excellent results. To evaluate the performance of ChatGPT on biomedical related tasks, this paper presents a comprehensive benchmark study on the use of ChatGPT for biomedical corpus, including article abstracts, clinical trials description, biomedical questions and so on. Through a series of experiments, we demonstrated the effectiveness and versatility of Chat-GPT in biomedical text understanding, reasoning and generation.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BERTMeSH: Deep Contextual Representation Learning for Large-scale High-performance MeSH Indexing with Full Text 96%
- Improving dictionary-based named entity recognition with deep learning 96%
- Lifestyle factors in the biomedical literature: An ontology and comprehensive resources for named entity recognition 95%
Similar papers in this journal
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 93%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 92%
Similar papers in this journal
Similar papers in this journal
- A Sequence Labeling Framework for Extracting Drug-Protein Relations from Biomedical Literature 96%
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 94%
- LSD600: the first corpus of biomedical abstracts annotated with lifestyle–disease relations 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.