Abstract
The classification of traffic accident causes from written description is of utmost importance for enhancing road safety and informing resource allocation strategies. Traditional approaches to accident classification rely on rigid rule-based systems and conventional machine learning models, which often struggle with unstructured accident descriptions and miss important contextual cues, ultimately limiting classification accuracy. These gaps limit context capture, lowering classification accuracy and reducing their usefulness for safety analysis. Therefore, this study aims to compare the performance of several advanced Large Language Models (LLMs)—ChatGPT-4, Claude-3 opus, Mistral small 3.2, DeepSeek-V3, and Gemini-1.5 pro in identifying the reasons behind traffic accidents. A dataset of 2,886 accident reports acquired from the Abu Dhabi Traffic Department was utilized, covering 37 accident cause categories. The models were given the accident descriptions and, using a zero‑shot learning approach, generated single‑label predictions. Their performance was then evaluated using accuracy, precision, recall, F1 score, unclassified rate, API error rate, and computation time. The experimental results revealed that all models demonstrated promising performance, with ChatGPT achieving the higher results (Accuracy: 0.7308; Precision: 0.7588), closely followed by Mistral (Accuracy: 0.7065; Precision: 0.7249). In terms of computational time, ChatGPT took slightly longer per prediction (0.086 s compared to Mistral’s 0.069 s), though both kept API error rates low. The misclassified cases were recorded due to common challenges: descriptions were often vague or incomplete, making it difficult for the models to predict the correct class. The findings provide practical guidance for selecting LLMs in traffic safety applications, balancing classification performance with efficiency and reliability considerations.
| Original language | English |
|---|---|
| Article number | 107231 |
| Journal | Safety Science |
| Volume | 200 |
| DOIs | |
| State | Published - Aug 2026 |
Keywords
- Accident cause classification
- Large language models
- Road safety
- Transportation safety
Fingerprint
Dive into the research topics of 'Decoding crash narratives: a comparative evaluation of large language models for accident cause classification'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver