Back

High-performing Multi-task Model of Urinary Tract Dilation (UTD) Classification for Neonatal Ultrasound Reports Through Natural Language Processing

Hua, Y.; Mukkamala, A.; Estrada, C.; Li, M. L.; Wang, H.-H.

2024-01-24 urology
10.1101/2024.01.23.24301680 medRxiv
Show abstract

ObjectiveThe urinary tract dilation (UTD) classification system provides objective assessment relevant to hydronephrosis management for children. However, the lack of uniform language regarding UTD in radiology reports leads to significant difficulty in both clinical management and research. We seek to develop a unified multi-task/multi-class model that can effectively extract UTD components and classifications from early postnatal ultrasound (US) reports. MethodsRadiology records from our institution were reviewed to identify infants aged 0-90 days undergoing early ultrasound for antenatal UTD. The report and images were reviewed by the study team to create the ground truth of UTD classification and components (primary outcome). Bio_ClinicalBERT, a variant of the Bidirectional Encoder Representations from Transformers (BERT) model, was used as the embedding layers of the classification model. The model was fine-tuned with 11 linear classification layers. All but the last BERT layer were frozen during the fine-tuning process. The model performance was evaluated with five-fold cross-validation with an 80:20 train-test ratio. Results2460 early (0-90 days) US reports were included. The five-fold cross-validated model performance is satisfactory (Weighted F1 > 0.9 for all UTD components). We report the weighted F1 scores, accuracies, and standard deviations for all 11 tasks and their average performance. ConclusionsBy applying deep state-of-the-art NLP neural networks, we developed a high-performing, efficient, and scalable solution to extract UTD components from unstructured ultrasound reports using one single multi-task model. This can potentially help standardize and facilitate large-scale computer vision research for pediatric hydronephrosis. Key Words: machine learning, efficiency, ambulatory care, forecasting

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Journal of Medical Imaging
11 papers in training set
Top 0.1%
19.3%
2
PLOS Digital Health
106 papers in training set
Top 0.1%
19.3%
3
Diagnostics
50 papers in training set
Top 0.1%
13.0%
50% of probability mass above
4
PLOS ONE
5266 papers in training set
Top 32%
4.5%
5
npj Digital Medicine
118 papers in training set
Top 1%
3.4%
6
Computers in Biology and Medicine
128 papers in training set
Top 1%
3.3%
7
Scientific Reports
3612 papers in training set
Top 38%
2.8%
8
Frontiers in Medicine
120 papers in training set
Top 1%
2.5%
9
Journal of the American Medical Informatics Association
71 papers in training set
Top 1%
2.5%
10
PLOS Computational Biology
1863 papers in training set
Top 13%
2.0%
11
IEEE Access
35 papers in training set
Top 0.7%
1.8%
12
Journal of Biomedical Informatics
47 papers in training set
Top 0.9%
1.4%
13
Biology Methods and Protocols
61 papers in training set
Top 1%
1.4%
14
European Radiology
15 papers in training set
Top 0.4%
1.2%
15
Communications Medicine
113 papers in training set
Top 3%
1.2%
16
BMC Medical Informatics and Decision Making
43 papers in training set
Top 1%
1.1%
17
JMIR Medical Informatics
18 papers in training set
Top 0.7%
1.1%
18
eBioMedicine
183 papers in training set
Top 5%
1.0%
19
International Journal of Medical Informatics
26 papers in training set
Top 1%
0.9%
20
JAMIA Open
42 papers in training set
Top 1%
0.9%
21
Modern Pathology
22 papers in training set
Top 0.4%
0.9%
22
Analytical Chemistry
218 papers in training set
Top 2%
0.9%
23
GigaScience
212 papers in training set
Top 5%
0.6%
24
Cancers
213 papers in training set
Top 5%
0.6%
25
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 1%
0.6%
26
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 2%
0.5%
27
Artificial Intelligence in Medicine
17 papers in training set
Top 0.9%
0.5%
28
BMJ Health & Care Informatics
15 papers in training set
Top 1%
0.5%