Classification of Animals with Different Deep Learning Models
Özkan İni̇k, Bülent Turan · DergiPark (Istanbul University) · 2018
The purpose ofthis study is that using different deep learning models for classification of14 different animals. Deep Learning, an area of artificial intelligence, hasbeen used in a wide range of recent years. Especially, it using in advancedlevel of image processing, voice recognition and natural language processingfields. One of the most important reasons for using a large field in imageanalysis is that it performs the feature extraction itself on the image andgives high accuracy results. It performs learning by creating at differentlevels representations for each image. Unlike other machine learning methods,there is no need of an expert for feature extraction on the images. ConvolutionNeural Network (CNN), which is the basic architecture of deep learning models, consistsof different layers. These are Convolution Layer, ReLu Layer, Pooling Layer andFull Connected Layer. Deep learning models are designed using different numbersof these layers. AlexNet and VggNet models are used for classified of 14different animals. These animals are Horse, Camel, Cow, Goat, Sheep, Wolf, Dog,Cat, Deer, Pig, Bear, Leopard, Elephant and Kangaroo respectively. Animals thatare most likely to encounter when during driving road were selected. Becausethinking this work to be a preliminary work for the control of autonomousvehicle driving. The images of animals are collected in color (RGB) on theinternet. In order to increase the data diversity, images were also taken fromthe ready data sets. A total of 150 images were collected with 125 training and25 test data for each animal. Two different data sets have been created, witheach image having dimensions of 224x224 and 227x227. As a result of the study,the classification of the animals was realized with %91.2 accuracy with VggNetand %67.65 with AlexNet. The high error rate in AlexNet is due to the smallnumber of layers in the network and the high selection of parameter values. Forexample, the filter size in the convolution layer in AlexNet architecture is11x11 and the number of stride is 4. This situation causes data loss intransferring the information to the next layer. In contrast, VggNet has afilter size of 3x3 and a number of steps of 1, there is no data loss in thetransfer to the next layer.