Exploiting ItalWordNet Taxonomies in a Question Classication Task
Francesca Bertagna, Linguistica Computazionale, Via G. Moruzzi · 2004
The paper presents a case-study about the exploitation of ItalWordNet for Question Answering. In particular, we will explore the access to ItalWordNet when trying to derive the information that is crucial for singling out the answers to Italian Wh-questions introduced by the interrogative elements Quale and Che. The paper describes some aspects arised during the rst phase of the work carried out for a Ph.D. research 1 dedicated to the exploration of the role of linguistic resources (from now on LRs) in a Question Answering (QA) application. The leading idea of the thesis is that the testing activity can highlight potentialities, together with problems and limitations, of the bulk of information collected during the last two decades by linguists and computational linguists. Altought LRs are not conceived to meet the requirements of a specic task (but rather to represent a sort of repository of information of general interest), they are signicant sources of knowledge that should allow systems to automatically perform inferences, retrieve information, summarize texts, translate words in context from a language to another etc.. Computational lexicons storing semantic information, in particular, are supposed to provide a description of the meaning of the lexical units they collect. It is interesting to evaluate what is the heuristic value of such description and to what extent it is exploitable and useful to perfom specic tasks (e.g. in matching question and answer). Tons of papers have been written about the use of WordNet in IR and in QA and the time is mature to test also resources dedicated to languages other than English, such as, for instance, the Italian component of the EuroWordNet project (i.e. ItalWordNet). The rst two sections of the paper will be devoted to briey introduce the IWN project and the preliminar steps for question analysis. The core of the paper is represented by a sort of case-study dedicate to the description of the way the QA system can access the semantic information in IWN with the goal to derive what we call the Question Focus, the information crucial to match question and answer. Unfortunately, we are not able to provide validated results yet. We are in the process of assembling the available components of the QA downstream (the search engine, the chunker and the dependency parser, as well as the LRs) and we hope to be able to provide the rst results soon. The current research is not collocated within a funded project but we hope to nd occasion of fundings in the future.