ToKSA - Tokenized Key Sentence Annotation - a Novel Method for Rapid Approximation of Ground Truth for Natural Language Processing
Fairfield, C. J.; Cambridge, W. A.; Cullen, L.; Drake, T. M.; Knight, S. R.; Masson, N.; Mills, N. L.; Pius, R.; Shaw, C. A.; Wu, H.; Wigmore, S. J.; Spiliopoulou, A.; Harrison, E. M.
Show abstract
ObjectiveIdentifying phenotypes and pathology from free text is an essential task for clinical work and research. Natural language processing (NLP) is a key tool for processing free text at scale. Developing and validating NLP models requires labelled data. Labels are generated through time-consuming and repetitive manual annotation and are hard to obtain for sensitive clinical data. The objective of this paper is to describe a novel approach for annotating radiology reports. Materials and MethodsWe implemented tokenized key sentence-specific annotation (ToKSA) for annotating clinical data. We demonstrate ToKSA using 180,050 abdominal ultrasound reports with labels generated for symptom status, gallstone status and cholecystectomy status. Firstly, individual sentences are grouped together into a term-frequency matrix. Annotation of key (i.e. the most frequently occurring) sentences is then used to generate labels for multiple reports simultaneously. We compared ToKSA-derived labels to those generated by annotating full reports. We used ToKSA-derived labels to train a document classifier using convolutional neural networks. We compared performance of the classifier to a separate classifier trained on labels based on the full reports. ResultsBy annotating only 2,000 frequent sentences, we were able to generate labels for symptom status for 70,000 reports (accuracy 98.4%), gallstone status for 85,177 reports (accuracy 99.2%) and cholecystectomy status for 85,177 reports (accuracy 100%). The accuracy of the document classifier trained on ToKSA labels was similar (0.1-1.1% more accurate) to the document classifier trained on full report labels. ConclusionToKSA offers an accurate and efficient method for annotating free text clinical data.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 93%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 92%
- Artificial Intelligence Model for Analyzing Colonic Endoscopy Images to Detect Changes Associated with Irritable Bowel Syndrome 92%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 93%
- Remote Covid Assessment in Primary Care (RECAP) risk prediction tool: derivation and real-world validation studies 90%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.