A Deep Learning Co-training Framework for e-book Classification
Tsui-Ping Chang, Hung-Ming Chen, Jianqun Chen · 2020
Automatic e-book classification is an important research issue since more and more people read and acquire information on their mobile devices (i.e., smartphones). Many writers digitize their works for users to acquire data on their mobile devices and result in the number of e-books has grown significantly. An e-book is a digital or electronic book that is formatted into a file that can be read on a mobile device. Some features of books are used in e-books. However, the important difference is that an e-book has a lot of images to describe the contents or knowledge of writers. Many researches proposed their methods of e-book classification to make users easily to find out and then read an e-book on their mobile devices. In these methods, an e-book can be categorized by several criteria. One of it is based on its type, i.e., novel, reference, and encyclopedia. Another is based on its topic, i.e., economy, religion, and technical. The classification systems based on the topics usually use well-known methodology such as Dewey Decimal Classification, in which, every category reflected by a decimal. These researches use Naïve Bayes Classifier and focus on automatic thesis classification. As deep learning proves its usefulness in an ever greater number of applications, there is a rise in demand for faster computational resources to train ever complex learning-based models. Based on the concept of deep learning, some researches proposed their methods to automatic e-book classification. W. A. Wiegand proposed a convolutional-neural-network (CNN) book label recognition algorithm to find out the misplaced books. On the other hand, four steps are illustrated in X. Yang et al. First, the keywords are extracted from the description data of e-books. Then, the description data is modeled as vectors of keywords. Third, the statistical categorization rules are obtained from meta-information of e-books. Finally, the vectors and the statistical categorization rules are combined to obtain an classification model. In this paper, a deep learning co-training framework (namely DLC) is proposed for improving the accuracy of automatic e-book classification. DLC combines the features of texts and images in e-books to co-training an e-book classifier. In order to increase the variety of feature sets, the texts in e-books are represented as vectors by Word2Vec. The images in e-books are translated and then combined into the vectors. Furthermore, DLC adopts the softmax regression function to co-train the combined features (i.e., vectors and images) in CNN to improve the accuracy of e-book classification. The experimental results demonstrate that our DLC has higher accuracy than other e-book classifiers.