Mask R-CNN with Multi-Backbones - A Comparative Analysis
Rufus Rubin, Chinnu Jacob, S. M. Anzar, Alavikunhu Panthakkan · 2022
Instance segmentation is a computer vision task for detecting and localizing an object in an image. It involves identifying, segmenting, and categorizing each object in an image. The idea behind instance segmentation is to categorize objects that belong to the same class into different instances. With instance segmentation, one can locate the bounding boxes of each instance as well as the object segmentation maps for each instance, allowing one to determine the number of instances in the image. Mask Region-based convolutional neural network is the cutting-edge method for instance segmentation. The backbone architecture of the mask R-CNN consists of a feature pyramid network, a region proposal network, and a region of interest alignment network. In this paper, three CNN models such as ResNet101, ResNet50, and MobileNetV1 are used as backbone network structures to compare the mask R-CNN architecture. The performance of these models is compared with the Penn- Fudan dataset using performance metrics such as mean Average Precision, mean Average Recall, and F1-score. The ResNet-101 output achieves a mean Average Precision of 98.3% over the other two models. The performance scores and losses of the three models were also discussed. Experimental studies show that ResNet-101 outperforms the other two models.