Recent Advances in Neural Networks Structural Risk Minimization Based on Multiobjective Complexity Control Algorithms
D.A.G. Vieira, J.A. Vasconcelos, Rodney Saldanh · Machine Learning · 2010
Nowadays, neural networks (NNs) are widely applied in the solution of several real world problems.They have been successfully used in many fields such as chemistry, physics, engineering, and bio-informatics among others.However, their use often relies on some handcrafted settings, such as the number of layers and neurons.This chapter will discuss the Structural Risk Minimization (SRM) problem using some multiobjective optimization concepts.Both are closely related to the classical Tikhonov's regularization scheme, and, it is also exploited in this work.A neural network is a learning machine capable to describe, to the input x, the set of functions F = { f (x, w) : x ∈ X, w ∈ W}, where W is the space of possible weights.Given a supervisor which defines an output vector y ∈ Y (desired output), for a given input x, according to the conditional distribution F(y|x), the ultimate goal in the learning problem is to find w ∈ W that best approximates the supervisor answer given some measure.To some loss function L(.), the expected risk (error) can be defined as, Vapnik (1998): R(w) = L (y, f (x, w)) dF(x, y).(variance dilemma, S. Geman & Doursat (1992).The expected mean-squared error between f (•) and the expected value of y given x, E[y|x], can be written as:where E T [.] is the expected value given a set T. The first term in the right hand side of ( 3) is known as bias, and the second one as variance.The variance term measures the sensibility of the approximating function given a data set T. To control the variance, models with less complexity should be generated, i.e., they cannot change too much to a given data T. On the other hand, some bias is inserted in the problem when the complexity is limited, thus, this should be controlled.This chapter is organized as follows.First, the regularization theory from Tikhonov (1963), a well-known technique to solve linear ill-posed problems, will be introduced together with the residual method from Phillips (1962) and the quasi-solutions from Ivanov (1962;1976).It is shown, using Singular Value Decomposition (SVD), the relationship between these methods and the Wiener's filter.After that, the Structural Risk Minimization (SRM), and the multiobjective learning will be discussed.These methods are closely related, and, some of their main aspects will be discussed.Inspired on the Tikhonov's regularization it will be discussed the well-known weight decay (WD) method for NNs, Hinton (1989).However, it will be clarified that this method is not consistent if the functions are not convex, which is usually the case.To overcome that, it is introduced the generalized Tikhonov's regularization based on a Q-norm for Parallel Layers Perceptrons (PLPs).Finally, some results are presented.Recent advances in Neural Networks Structural Risk Minimization based on multiobjective complexity control algorithms 105 y z Transmitter Receiver