AMICA: the AT&t mixed initiative conversational architecture.

Roberto Pieraccini, Esther Levin, Wieland Eckert · 1997

In this paper we show how it is possible to design and implement a general architecture that is suitable for the rapid development of human/machine natural language, mixed initiative dialogue systems. The architecture proposed here relies on the assumption that a dialogue system can be modularized into different actions or functions that can be designed separately and implement basic aspects of the dialogue behavior, and a strategy that is fairly independent of the particular application. INTRODUCTION Developers of human/machine natural language dialogue systems often state that one of the main problems of the field is that of finding a general framework that would easily fit different applications, and would allow for a rapid development of new system. One of the reasons of this difficulty lies in the lack of separation, in many existing systems, among the different levels of competence that intervene during the dialogue activity. When one looks at the dialogue as the result of logical activity whose basic rules are independent of the particular application, the language, and the medium in question (as for instance in [5]), the design of dialogue systems becomes more of an engineering problem and less of an art. For instance, in the design of a form filling application the dialogue flow is generally represented by a tree that takes into account all the possible outcomes,. Instead one could design a function that implements the basic logic principle that, when some pieces of information is not present in the current memory, the best move for the system is asking for it. In this spirit we present AMICA, a general model of a mixed initiative dialogue system based on the identification of a set of general, logically motivated functions, called dialogue actions. We think that for certain classes of dialogues there exists a finite (and small) number of such actions that can be implemented in a general way and parametrized in order to be used for different applications. In this work we restrict the discussion to those dialogue systems that are devoted to the extraction of information from a database. For an effective dialogue the machine needs to be able to accomplish the actions in the following inventory: Understanding: that is the transduction of the user input (generally written or spoken natural language) into a formal representation conveying the semantics of the message. Verbalization: that consists in transducing the machine output into a form that is promptly understood by the user (e.g. natural language). We distinguish between data verbalization, that requires specific knowledge about the semantic structure of the domain, and sentence verbalization that needs general, domain independent, knowledge. Contextual Interpretation: consists in the interpretation of the current user input in terms of the history of the interaction. It includes both dialogue interpretation, namely the ability of dealing with expectations set by the machine itself (e.g. disambiguation on the basis of previous questions asked by the machine), and discourse interpretation, namely the ability of taking into account the context set by the user during the course of the whole interaction. In both cases it is necessary to be able to recognize ambiguity and to formulate disambiguation questions. Constraint Consistency Verification: consists in the operation of spotting the presence of inconsistent sets of constraints provided by the user. These inconsistencies might be the result of user errors, like false presupposition, or simply errors of the recognizer. Data Retrieval: consists in the assembly and submission of a query to a local or remote database according to the current information gathered by the system. Constraining: is the operation the system performs when asking the user for additional information. This operation is required due to both the limited bandwidth of the communication protocol, and the limited capacity of humans to analyze big sets of data. In general constraining is required when an under-constrained query produces too many results. In certain cases it is possible to predict that a query is under-constrained without actually accessing the database, for instance by specifying a set of minimal constraints for each particular topic. Relaxation: consists in the ability of analyzing the failure of a database query (i.e. empty set of data), and proposing the user alternative solutions obtained by relaxing one or more constraints. Sequencing: is the operation required when presenting a long list of items that exceed the capability of the bandwidth and cannot be further reduced by constraining. Sequencing should allow the user to navigate through the list. Although there are other basic functions the dialogue system should be able to perform, like for instance ambiguity resolution and the capability of coping with possible noisy input (e.g. the errors of a speech recognizer), we limit here our discussion to the previous functions. Once these (and maybe other) basic functions of dialogue have been defined we need two other components in order to build a dialogue system, namely a representation of the current status of knowledge of the machine (called state [6]), and a mechanism that invokes the required function when needed, that we call strategy. In our implementation the strategy is represented by recursive transition network, whose arcs represent conditions on the state, and whose nodes represent handles to the above mentioned functions. The idea of implementing the dialogue as a constraining/relaxation activity can be found in [2], in [4] most of the functions described here were also and in [5] the idea of separating the dialogue activity into several levels of competence, starting with the innermost more general logical functions is intro-

Read the paper · More papers on PaperTik