Learning by α-divergence
Ryotaro Kamimura, Shinobu Nakanishi · 2002
In the present paper, we propose a new cost function, called /spl alpha/-divergence, which is a generalized version of the relative entropy or the Kullback's divergence measure in neural network. The most fundamental characteristics of this /spl alpha/-divergence are summarized by the following three points: 1) by changing the parameter /spl alpha/ for the /spl alpha/-divergence, multiple cost functions can be obtained to be used for different purposes or problems; 2) /spl alpha/-divergence is effective in direct proportion to the error between targets and outputs, eliminating the derivative of the sigmoidal function; and 3) the /spl alpha/-divergence has an effect on eliminating saturated units. We formulated an update rule to minimize /spl alpha/-divergence, and applied the method to the acquisition of the grammatical competence. Experimental results confirmed marked improvement in the generalization by using /spl alpha/-divergence. This improvement is due to the property of /spl alpha/-divergence whose derivative is effective especially for eliminating saturated units.>