Ensembles of diverse neural networks

M. A. H. Akhand · University of Fukui Library (University of Fukui) · 2009

An ensemble with several neural networks has become very popular to achieve \tbetter performance over a single neural network. Several neural networks are trained \tand combined their decisions for the ensemble’s decision. An important condition \tassociates in designing the ensemble is proper diversity among component networks so \tthat the failure of one network can be compensated by others. Since function to \tapproximate by a network is learned from its training data, data sampling (i.e., different \ttraining sets for different networks) is considered as an effective approach to produce \tdiversity among component networks. \tSeveral data sampling based existing ensemble methods are investigated in this \tthesis. A large number of benchmark problems from different application domains were \tconsidered for evaluating the methods. Experimental results reveal that no single \tmethod is superior to others for all the problems; a particular method is shown better \tthan others for a subset of problems. Therefore, to choose an ensemble method for a \tparticular problem, this study discussed proficiency of the methods based on the \tproblem features. However, among the methods, bagging and AdaBoost were found \tbetter than others. Both the methods explicitly create different training sets for different \tnetworks using bootstrap sampling technique. Negative correlation learning, on the \tother hand, interactively trains all the networks with same original training data. Due to \ttraining time interaction, it implicitly motivates different networks toward different \ttraining subspaces and is found competitive with bagging and AdaBoost. Finally, study \tof existing ensemble methods leads to develop better ensemble construction method. \tIn an ensemble, component networks solve a given problem individually and \tcombine their decisions for ensemble’s decision. Therefore, it is important to construct an ensemble with appropriate networks. Existing ensemble methods, in general, train a \tpredefined number of networks and consider all of them for final ensemble. Therefore, \tperformance of an ensemble might be poor when any network performs very badly. On \tthe other hand, ensemble with properly selected networks might always perform better. \tIn this context, an ensemble method is proposed in this thesis that first creates a pool of \tdiverse networks and then apply selection scheme to construct an ensemble with \tappropriate networks. Three data sampling techniques are considered to create network \tpool, and two selection schemes are investigated. Based on experimental results on a \tlarge number of benchmark problems, the proposed method is found better than other \ttraditional methods with concise ensemble. \tIn this thesis, an ensemble method is also investigated that automatically determines \ta minimal ensemble architecture for a given problem. To determine minimal \tarchitecture, it starts with a single network with a minimal number of hidden units. \tDuring training process, it adds additional network(s) with cumulative number(s) of \thidden units and the added network specializes in the previously unsolved portion of the \tinput space. Finally all the networks are trained simultaneously to improve the \tgeneralization ability. When a single network is shown to achieve acceptable result for a \tproblem, it does not build and returns the single network for the problem. The proposed \tmethod is found competitive with existing ensemble methods with minimal architecture \twhen tested on benchmark problems. \tFurther, an indirect communication scheme among the networks is investigated in \tthis study when they are trained for an ensemble and proposed progressive interactive \ttraining scheme. In the proposed scheme, networks are trained one after another and \tinteraction among the networks is maintained indirectly via an intermediate space called \tinformation center. The idea of using indirect communication is conceived from the \tcommunication among biological ants via pheromone. An individual ant decides its \ttravelling path based on existing pheromone on the trail and also it deposits pheromone \ton its travelling path. The indirect interaction has several benefits over direct interaction \tscheme of negative correlation learning. The experimental results show that ensembles \tconstruction with the proposed training scheme performs well. The idea of indirect \tinteraction is also incorporated with bagging and AdaBoost, and is shown to improve \ttheir performance.

Read the paper · More papers on PaperTik