A thangka classification study based on Omni-Dimensional Dynamic Convolution and Resnet

Zhongjie Tang, Chunhua Pan, Xuegang Zhang · 2024

Thangka is a unique art form in Tibetan culture. Due to its complex texture features and diverse types, the traditional image classification method cannot achieve rapid and accurate classification of Thangka images, which is of great significance for people to study, protect and appreciate Thangka images. To solve this problem, this paper proposes a full-dimensional dynamic convolution module based on OMNI-DIMENSIONAL DYNAMIC CONVOLUTION (ODConv) and ResNet residual neural network to form a full-dimensional dynamic residual neural network model. The network model is used to classify Thangka images. ODConv module uses a convolution method that pays attention to the dynamics of spatial space, input channel, output channel and other dimensions, which is called full-dimensional dynamic convolution. Therefore, ODConv module can greatly improve the feature extraction capability of convolution. The purpose of this study is to explore the application effect of ODconv module in Thangka image classification task after it is added to ResNet network. The ODConv module is added to the ResNet101 residual neural network of layer 101, and the network model after the addition of the module is the ODConv-Resnet101 full-dimensional dynamic convolutional residual network model of layer 101. The improved model and the original model are classified and tested on the Thangka image data set, and finally the confusion matrix is obtained. The accuracy and recall rate of the confusion matrix are used to judge the model. The experiment shows that ODConv-ResNet101 network model has the best classification effect, and the accuracy rate and recall rate are 89.25% and 88.87% respectively, while the accuracy rate and recall rate of ordinary Resnet101 network model are only 82.51% and 81.69%. This indicates that after the addition of ODConv module, ResNet residual neural network model has improved its ability to extract features and classify images, and can perform thangka image classification tasks more efficiently and accurately.

Read the paper · More papers on PaperTik