Image Captioning for Visual Surveillance System
Ade Nurhopipah, Siti Alvi Sholikhatin, Anan Widianto · 2023
Closed Circuit Television (CCTV) for surveillance has evolved from a passive manual system to an automated, integrated intelligent system. However, image processing in this system is generally based on detection and classification, so the resulting output does not give as good an impression as the human language description. In this research, we implemented Image Captioning with Bahasa output on office area CCTV data. This model used pre-trained EfficientNetV2s as the image extractor and GRU as the language generator. We explored the best architecture by adding attention blocks, dropout, and batch normalization layers. We also searched for the best hyperparameter values for the optimizer, activation function, weight initialization, and batch size. The best model evaluated and produced a moderate average score for BLEU-1 of 0.5710 and a Meteor score of 0.3733. In general, the output description is acceptable. However, there are still many captions that have correct grammar but have semantic errors. In the future, it is necessary to develop a more reliable model to be implemented as a new generation monitoring system that produces output in a more varied and natural language.