A Cross-Domain Threat Screening and Localization Framework Using Vision Transformers and Self-supervised Learning

Ammara Nasim, Muhammad Usman Akram, Asad Mansoor Khan, Muhammad Belal Afsar Khan, Taimur Hassan · 2024

Due to the ever-changing global security landscape, various countries are constantly updating their safety protocols at designated airports and train stations. Consequently, there is an expanding inventory of prohibited items that are not permitted to be carried in luggage. This renders training of prior models on large datasets but lesser threat classes, ineffective. The scarcity of extensive, accurately labeled X-ray images with new threat items hinders the feasibility of training data-intensive, fully supervised deep models, primarily due to the high cost and time associated with manual labeling. In this paper, a self-supervised threat detection approach is discussed that is combined with vision transformer(VIT)-based pre-screening. The approach is adaptable to new classes and can easily detect new threat classes in the inference stage using very few annotated support images. The VIT based classification module showcases promising results with an accuracy of 98% and F1-score of 99%. The self-supervised threat detection module surpasses other self-supervised and weakly supervised frameworks with Mean Average Precision (mAP) of 0.89 and an Intersection over Union (IoU) value of 0.70 on GDXRAY dataset. The framework performs exceptionally well on SIXRAY dataset with an IOU score of 0.67 and mAP value of 0.865 despite being trained on GDXRAY dataset.

Read the paper · More papers on PaperTik