Back

A Natural Language Processing Tool to Extract Quantitative Smoking Status from Clinical Narratives

Yang, X.; Yang, H.; Lyu, T.; Yang, S.; Guo, Y.; Bian, J.; Xu, H.; Wu, Y.

2020-11-04 health informatics
10.1101/2020.10.30.20223511 medRxiv
Show abstract

This study presents a natural language processing (NLP) tool to extract quantitative smoking information (e.g., Pack-Year, Quit Year, Smoking Year, and Pack per Day) from clinical notes and standardized them into Pack-Year unit. We annotated a corpus of 200 clinical notes from patients who had low-dose CT imaging procedures for lung cancer screening and developed an NLP system using a two-layer rule-engine structure. We divided the 200 notes into a training set and a test set and developed the NLP system only using the training set. The experimental results on the test set showed that our NLP system achieved the best F1 scores of 0.963 and 0.946 for lenient and strict evaluation, respectively. NoteAccepted as a presentation at the 2020 IEEE International Conference on Healthcare Informatics (ICHI) Workshop on Health Natural Language Processing (HealthNLP 2020). https://ohnlp.github.io/HealthNLP2020/healthnlp2020#.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.