Improving the CNN Model for Arabic Crime Tweet Detection Based on an Intelligent Dictionary
Zainab Khyioon Abdalrdha, Abbas M. Al-Bakry, Alaa Kadhim Farhan · 2023
Rising crime rates in Iraq and the Arab World pose significant challenges for law enforcement, exacerbated by vast data on criminal activity, technological advances, and high population density. This research aims to extract credible information from this data to understand crime better and assist in future prevention efforts. One challenge is the complexity of the Arabic language in crime detection. The study introduces an innovative approach using deep learning techniques, including creating an intelligent dictionary for detecting crime-related content in Arabic tweets. This method employs the Aho-Corasick algorithm to address the dynamic nature of language and evolving criminal terminology. The methodology applies deep learning and machine learning techniques, emphasizing natural language processing (NLP) for preprocessing and extracting relevant tweet characteristics. Created a labeled tweet dataset and classified it into various crime-related categories. Deep learning models, such as an improved CNN model based on chi-square testing and Grid search, were trained and evaluated using precision, recall, F1-score, accuracy, and MAP accuracy metrics. The analysis of an Arabic tweet dataset, comprising 18,493 tweets with ten features, demonstrated that the improved CNN model achieved a high accuracy of 99.25% and a macro F1-score of 99.9%. These results indicate that the proposed deep learning model outperforms others regarding accuracy.