High-performing Multi-task Model of Urinary Tract Dilation (UTD) Classification for Neonatal Ultrasound Reports Through Natural Language Processing
Hua, Y.; Mukkamala, A.; Estrada, C.; Li, M. L.; Wang, H.-H.
Show abstract
ObjectiveThe urinary tract dilation (UTD) classification system provides objective assessment relevant to hydronephrosis management for children. However, the lack of uniform language regarding UTD in radiology reports leads to significant difficulty in both clinical management and research. We seek to develop a unified multi-task/multi-class model that can effectively extract UTD components and classifications from early postnatal ultrasound (US) reports. MethodsRadiology records from our institution were reviewed to identify infants aged 0-90 days undergoing early ultrasound for antenatal UTD. The report and images were reviewed by the study team to create the ground truth of UTD classification and components (primary outcome). Bio_ClinicalBERT, a variant of the Bidirectional Encoder Representations from Transformers (BERT) model, was used as the embedding layers of the classification model. The model was fine-tuned with 11 linear classification layers. All but the last BERT layer were frozen during the fine-tuning process. The model performance was evaluated with five-fold cross-validation with an 80:20 train-test ratio. Results2460 early (0-90 days) US reports were included. The five-fold cross-validated model performance is satisfactory (Weighted F1 > 0.9 for all UTD components). We report the weighted F1 scores, accuracies, and standard deviations for all 11 tasks and their average performance. ConclusionsBy applying deep state-of-the-art NLP neural networks, we developed a high-performing, efficient, and scalable solution to extract UTD components from unstructured ultrasound reports using one single multi-task model. This can potentially help standardize and facilitate large-scale computer vision research for pediatric hydronephrosis. Key Words: machine learning, efficiency, ambulatory care, forecasting
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The Effect of Image Resolution on Automated Classification of Chest X-rays 93%
- Predicting Primary Site of Secondary Liver Cancer with a Neural Estimator of Metastatic Origin (NEMO) 92%
- A 3D CNN Classification Model for Accurate Diagnosis of Coronavirus Disease 2019 using Computed Tomography Images 92%
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 95%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 94%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 94%
Similar papers in this journal
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 91%
- Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images using transfer learning 90%
- An Explainable Web-Based Diagnostic System for Alzheimer's Disease Using XRAI and Deep Learning on Brain MRI 90%
Similar papers in this journal
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 94%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 94%
- Predicting semantic segmentation quality in laryngeal endoscopy images 93%
Similar papers in this journal
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 94%
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 93%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.