KGL: A Language for Learning
Kenneth A. Kaufman, Ryszard S. Michalski · 1997
In real-life data mining endeavors, the extraction of important knowledge may require many trials and errors, and multiple executions of different sequences of data mining operations. Such applications may pose a variety of tasks, for example, determining a characteristic description or a discriminant description of given classes of entities, optimizing an initial hypothesis according to a cost function, determining the most relevant attributes for a given task, selecting the most representative examples from a large example set, conceptually clustering cases into classes, predicting the class membership of a new example, generating a decision structure, automatically determining a learning curve, etc. The application of these sequences of programs can be time-consuming, laborious, and error prone. In response to these challenges, it is important for machine learning to develop a methodology for integrating diverse learning strategies so that a learning system can pursue different learning tasks and acquire different kinds of knowledge, depending on the problem at hand. We have developed a high-level language, called KGL, in which a data analyst can plan data-mining and knowledge discovery experiments using various operators for performing learning and discovery tasks. The presented approach implements a range of learning and knowledge processing programs as KGL operators. Using KGL, a user can specify a plan for applying these operators in a flexible and interdependent manner in pursuit of a desirable solution. The methodology and language are illustrated by a problem of detecting demographic and economic patterns from a database of 190 countries. The results show a great potential of the proposed approach for data mining.