Self-supervised visual learning for analyzing firearms trafficking activities on the Web

Sotirios Konstantakos, Despina Ioanna Chalkiadaki, Ioannis Mademlis, Adamantia Anna Rebolledo Chrysochoou, Georgios Th. Papadopoulos · 2023

Automated visual firearms classification from RGB images is an important real-world task with applications in public space security, intelligence gathering and law enforcement investigations. When applied to images massively crawled from the World Wide Web (including social media and dark Web sites), it can serve as an important component of systems that attempt to identify criminal firearms trafficking networks, by analyzing Big Data from open-source intelligence. Deep Neural Networks (DNN) are the state-of-the-art methodology for achieving this, with Convolutional Neural Networks (CNN) being typically employed. The common transfer learning approach consists of pretraining on a large-scale, generic annotated dataset for whole-image classification, such as ImageNet-1k, and then finetuning the DNN on a smaller, annotated, task-specific, downstream dataset for visual firearms classification. Neither Visual Transformer (ViT) neural architectures nor Self-Supervised Learning (SSL) approaches have been so far evaluated on this critical task. SSL essentially consists of replacing the traditional supervised pretraining objective with an unsupervised pretext task that does not require ground-truth labels. This paper evaluates a common CNN and a typical ViT architecture in combination with different SSL methods, comparing them against each other and against supervised pretraining in terms of downstream classification accuracy. Additionally, “CrawledFirearmsRGB” is introduced as a new, challenging image dataset for visual classification of firearms and other concepts related to on-line criminal networks. Finally, a new mixed pretraining objective is formulated that combines SSL and whole-image classification, under a multitask learning setting. The experimental results, indicate the superiority of certain SSL pretraining methods that cooperate well with the ViT architecture, even when the pretraining dataset is of a scale similar or identical to that of CrawledFirearmsRGB, despite the fact that no ground-truth labels are exploited.

Read the paper · More papers on PaperTik