Extracting traffic crash information from social media: an LLM-based approach

Reza Golshan Khavas, Mohammadreza Nosrati · Transportation Letters · 2026

Fast detection and analysis of traffic crashes are important steps toward improving road safety. While traditional data sources remain useful, social media platforms have become rich and rapidly growing channels for tracking such incidents. Still, the informal and unstructured nature of user-generated content makes accurate information extraction a challenge. This study explores the potential of using social media data mining and Large Language Models (LLMs) to automatically detect traffic crashes and extract relevant details from social media messages. The research was centered on Damavand County, a high-traffic area located on major routes near Tehran. A full framework was developed, including data collection, preprocessing, and fine-tuning of BERT-based models to classify messages into crash and non-crash categories and further categorize them into ten crash types, reaching accuracies of 91.1% and 89.7%, respectively. Additionally, a LLaMA 3.1 model was fine-tuned for a question-answering task focused on extracting crash location and casualty numbers. It achieved 97% accuracy in identifying fatalities and injuries, and BLEU and METEOR scores of 0.697 and 0.864 in location extraction. Comparison with official records showed that social media data identified 64.1% of crash hotspots found in official reports. Moreover, crashes with higher severity and specific types, like rollover or fall crashes, were more frequently reported. These findings highlight the practical value of social media as a supplementary data source and show how LLMs can facilitate this process of extracting useful crash-related information from social media.

Read the paper · More papers on PaperTik