Learning Novel Concepts in the Kinship Domain
Daniel M. Roy · 2004
This paper addresses the role that novel concepts play in learning good theories. To concretize the discussion, I use Hinton’s kinship dataset as motivation throughout the paper. The standpoint taken in this paper is that the most compact theory that describes a set of examples is the preferred theory—an explicit Occam’s Razor. The kinship dataset is a good test-bed for thinking about relational concept learning because it contains interesting patterns that will undoubtedly be part of a compact theory describing the examples. To begin with, I describe a very simple computational level theory for inductive theory learning in first-order logic that precisely states that the most compact theory is preferred. In addition, I illustrate the obvious result that predicate invention is a necessary part of any system striving for compact theories. I present derivations within the Inductive Logic Programming (ILP) framework that show how the intuitive theories of family trees can be learned. These results suggest that encoding regular equivalence directly into the training sets of ILP systems can improve learning performance. To investigate theories resulting from optimization, I devise an algorithm that works with a very strict language bias allowing all consistent rules to be entertained and explicitly optimized over for small datasets. The algorithm, which can be viewed as a special case implementation of ILP, is capable of learning a theory of kinship comparable in compactness to the intuitive theories humans use regularly. However, this alternative approach falls short as it is incapable of inventing the unary predicate sex to learn a more compact theory. Finally, I comment on the philosophical position of extreme nativism in light of the ability of these systems to invent primitive concepts not present in the training data.