Application of Large Language Models (LLM) for Automatic Classification of Work Accident Text Data: Verification of Accuracy and Practicality
Ando, H.; Matsugaki, R.; Yamakawa, S.; Ogami, A.
Show abstract
BackgroundFalls are the most frequent type of occupational accident, making the development of effective countermeasures an urgent issue. Traditional accident analysis relies on manual classification of text data by experts, a process that is both time-consuming and labor-intensive. Large Language Models (LLMs) offer the potential to significantly streamline this analysis process without the need for task-specific pre-training. ObjectiveThis study aims to automatically classify text data from occupational accidents using LLMs and to verify the accuracy and practicality of this approach. MethodsThe analysis targeted 2,619 fall-related injury cases in the health/hygiene and retail sectors, extracted from the 2021 Survey on Industrial Accidents database. The results of manual classification performed by experts in a previous study were used as the "ground truth." These were compared against the automatic classification results from four different LLMs (GPT-4.1, GPT-4.1 mini, GPT-4o mini, and o4-mini). Evaluation metrics included accuracy, precision, recall, F1-score, and Cohens kappa coefficient. The processing was conducted using OpenAIs Batch API, with processing time and costs also being measured. ResultsNewer generation models demonstrated a high rate of agreement with expert classifications across most categories, with the exception of "causal substance," generally achieving a Cohens kappa coefficient above 0.7. For the "accident location (indoor/outdoor)" category, the accuracy reached over 91%. Even for "causal substance," the category with the lowest accuracy, the reasoning model o4-mini achieved a kappa coefficient of 0.662. In terms of practicality, even when using the highest-performing model (o4-mini), the entire dataset was processed in approximately 90 minutes at a cost of about $11, demonstrating high cost-performance. ConclusionThis study demonstrates that LLMs can classify occupational accident text data with an accuracy comparable to manual expert analysis, but at a lower cost and higher speed. This method is expected to facilitate large-scale accident analysis, which has been challenging in the past, and contribute to the rapid development of evidence-based preventive measures for occupational accidents.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 93%
- Using a Bayesian network to classify time to return to sport based on football injury epidemiological data 93%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 92%
Similar papers in this journal
Similar papers in this journal
- Predicting Car Accident Severity in Northwest Ethiopia: A Machine Learning Approach Leveraging Driver, Environmental, and Road Conditions 93%
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 92%
- Early risk assessment for COVID-19 patients from emergency department data using machine learning 92%
Similar papers in this journal
- Accuracy of US CDC COVID-19 Forecasting Models 91%
- Machine Learning Based Clinical Decision Support System for Early COVID-19 Mortality Prediction 91%
- Applying machine-learning to rapidly analyse large qualitative text datasets to inform the COVID-19 pandemic response: Comparing human and machine-assisted topic analysis techniques 90%
Similar papers in this journal
- Quantifying Device Type and Handedness Biases in a Remote Parkinson’s Disease AI-Powered Assessment 92%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 91%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.