Multilayer perceptrons may learn simple rules quickly

Robert Urbanczik · Physical review. E, Statistical physics, plasmas, fluids, and related interdisciplinary topics · 1998

Zero-temperature Gibbs learning is considered for a connected committee machine with $K$ hidden units. For large $K,$ the scale of the learning curve strongly depends on the target rule. When learning a perceptron, the sample size $P$ needed for optimal generalization scales so that $N\ensuremath{\ll}P\ensuremath{\ll}KN,$ where $N$ is the dimension of the input. This holds even for a noisy perceptron rule if a new input is classified by the majority vote of all students in the version space. When learning a committee machine with $M$ hidden units, $1\ensuremath{\ll}M\ensuremath{\ll}K,$ optimal generalization requires $\sqrt{\mathrm{MK}}N\ensuremath{\ll}P.$

Read the paper · More papers on PaperTik