Active learning based on a hybrid neural network modeller

Yong-Long Wang · Abertay Research Portal (Abertay University) · 1998

Various m ethods are investigated for selecting training data for the purpose o f training neural networks.A new m ethod called M IQ R (M aximum Inter-Quartile R ange) is proposed for effectively selecting a concise set o f training data.In addition, the en sem b le concept is introduced in this new method.D ata selection is not unduly influenced by "o u tlie r s ", rather, it is principally dependent upon the "m a in s tr e a m " output o f the ensem ble networks.Encom passed in the new method is a very simple ancient Chinese philosophical idea, i.e. "the m in o rity o b e y s th e m a jo rity ".These techniques are nonparametric in the sense that several different neural networks com prise an ensem ble or com m ittee and co-operatively w ork together w ith each other to achieve a com m on goal.B ecause these are different neural networks (hybrid m odel), they can be complementary in the entire learning system, and therefore effectively enhance the entire learning system 's efficiency and accuracy.For learning, the neural networks attempt to actively select the m ost informative and important training data.The m ethods described in this thesis pleasingly satisfy this need, and com pare favourably with contending m ethods.M any experiments have been done to corroborate theoretical and empirical conjectures.The results are quite pleasing in that this new method is not only as "active learning" much better than "passive learning" both in data selection and in generalisation performance, but also outperforms other existing contending active learning m ethods.In particular, the results are very satisfying and interesting w hen the method is applied to discontinuous functions.Although the experiments are conducted with clean data selection, it should be easy to extend them to noisy data selection since the method developed is validated using unlabelled data.The algorithm developed for these m ethods has been rigorously tested, and proves to be highly autonom ous and robust.The m ethods developed here are not restricted to use on neural networks.M ore generally, they can be applied to other scientific research and econom ic fields, even educational and sociological behaviour.7. L a b e lle d data a n d u n la b e lle d d a ta : In the mappings o f X Y (here X and Y can be any dimensional space), if X is known, then Y is also determined by Y = / (X ).Such data is referred to as labelled data; Otherwise, no know ledge is required o f the target function to be approximated, it is referred to as unlabelled data.v 8. N etw ork com plexity: the degree o f com plexity o f the network architecture is normally the number o f hidden units.9. N etwork learning efficiency, defined as the ratio o f the number o f data points resampled to the number o f w eights o f the network, using a m ethod for selecting examples; 10.Nonparametric'.parametric statistical procedures are dependent upon rigid assum ptions, for example, that samples have been drawn from normally distributed populations with equal variance, or tests based on Student's t distribution.In contrast, nonparametric procedures are not concerned with population parameters or sample distributions, being valid under very general assumptions.11.O ver-fit: A network that is not sufficiently com plex can fail to detect fully the signal in a com plicated data set, leading to under-fit.A network that is too com plex may fit the noise, not just the signal, leading to over-fit.Over-fit is especially dangerous because it can easily lead to predictions that are far beyond the range o f the training data with many o f the com m on types o f neural networks.Over-fit can also produce wild predictions in multilayer perceptrons even with noise-free data.12. R andom sam ple: data is selected in such a w ay that all points have an equal probability o f being chosen.13.Resam pling: the process o f selecting a training data set (no matter whether using an active or passive learning m ethod) is called resampling.14.R esam pling efficiency is defined as the ratio o f the number o f data points resampled to the total training iterations, and is a measure o f the training procedure's efficiency.vi

Read the paper · More papers on PaperTik