On-Line training with guide data: Shall we select the guide data randomly or based on cluster centers?

Yuya Kaneda, Qiangfu Zhao, Yong Liu · 2016

To retrain an existing multilayer perceptron (MLP) on-line using newly observed data, it is necessary to incorporate the new information while preserving the performance of the network. This is known as the “plasticitystability” problem. For this purpose, we proposed an algorithm for on-line training with guide data (OLTA-GD). OLTA-GD is good for implementation in portable/wearable computing devices (P/WCDs) because of its low computational cost, and can make us more independent of the internet. Results obtained so far show that, in most cases, OLTA-GD can improve an MLP steadily. One question in using OLTA-GD is how we can select the guide data more efficiently. In this paper, we investigate two methods for guide data selection. The first one is to select the guide data randomly from a candidate data set G, and the other is to cluster G first, and select the guide data from G based on the cluster centers. Results show that the two methods do not have significant difference in the sense that both of them can preserve the performance of the MLP well. However, if we consider the risk of “instantaneous performance degradation”, random selection is not recommended. In other words, cluster center-based selection can provide more reliable results for the user during on-line training.

Read the paper · More papers on PaperTik