On-line versus off-line learning in the linear perceptron: A comparative study

Osame Kinouchi, Nestor Caticha · Physical review. E, Statistical physics, plasmas, fluids, and related interdisciplinary topics · 1995

The spherical perceptron with N inputs and a linear output does not present optimal generalization if trained by minimization of the standard quadratic cost function E=1/2 ${\mathcal{J}}_{\mathrm{\ensuremath{\mu}}=1}^{\mathrm{\ensuremath{\alpha}}\mathit{N}}$ (${\mathit{b}}_{\mathrm{\ensuremath{\mu}}}$-${\mathit{h}}_{\mathrm{\ensuremath{\mu}}}$${)}^{2}$, where ${\mathit{b}}_{\mathrm{\ensuremath{\mu}}}$ and ${\mathit{h}}_{\mathrm{\ensuremath{\mu}}}$ are the outputs from the rule (teacher) and hypothesis (student) networks for the example \ensuremath{\mu} and there are \ensuremath{\alpha}N examples. We derive an optimal algorithm for on-line learning of examples which outperforms the iterative (off-line) standard algorithm for \ensuremath{\alpha} up to 0.71. The on-line optimized algorithm suggests a class of cost functions for off-line learning, which we then proceed to study using the replica method. The optimized cost function within that class has the suggestive form E=\ensuremath{\alpha}N[\ensuremath{\Gamma}(1/\ensuremath{\alpha}N) ${\mathcal{J}}_{\mathrm{\ensuremath{\mu}}=1}^{\mathrm{\ensuremath{\alpha}}\mathit{N}}$ [-lnP(${\mathit{b}}_{\mathrm{\ensuremath{\mu}}}$\ensuremath{\Vert}${\mathit{h}}_{\mathrm{\ensuremath{\mu}}}$)]-\ensuremath{\Gamma} lnZ], where Z is a normalization constant, P(${\mathit{b}}_{\mathrm{\ensuremath{\mu}}}$\ensuremath{\Vert}${\mathit{h}}_{\mathrm{\ensuremath{\mu}}}$) is the conditional probability of the output data ${\mathit{b}}_{\mathrm{\ensuremath{\mu}}}$ given the hypothesis output ${\mathit{h}}_{\mathrm{\ensuremath{\mu}}}$, and \ensuremath{\Gamma} is a learning parameter analogous to a temperature which decreases in a well defined manner along the learning process.

Read the paper · More papers on PaperTik