Image classification and object detection complexity optimization: Exploring deep learning models on camera trap and surveillance clips
Hayder Abdulameer Yousif, Zahraa Al‐Milaji · Results in Control and Optimization · 2026
Input image size for convolutional neural networks (CNNs) has played a major role in classification accuracy and network speed. Designing a large depth, scale, and resolution CNN model can not guarantee the best performance because of the problems of overfitting and memorization. On the other hand, object detection models have produced very low performance on event-triggered camera-trap images due to highly dynamic scenes. In this paper, we propose a framework for optimizing image classification in terms of performance and complexity by selecting the convenient deep learning model for each image. Based on the image sequence activation maps, we propose Resolution Selection Model (RSM) that generates a weight value for each image in the sequence. We utilize support vector machine (SVM) and the generated weight from RSM to select the appropriate deep learning model. We utilized EfficientNet models that have different input image resolutions to classify and detect the objects from the scaled images. Our results on camera-trap and surveillance images show the efficacy of the proposed method compared to the state-of-the-art architectures in terms of accuracy and computational complexity.