Lexicon and Corpus: a Multi-faceted Interaction
Nicoletta Calzolari · 1996
All Language Engineering (LE) applications require knowledge about words. The first basic operation that any NLP system must perform consists in recognizing the words of the input-text. This means i) to search for the corresponding entry in a Computational Lexicon for each word (or multi-word) of a text, and ii) to associate the linguistic information provided by the lexical entry and relevant for that particular application, to the word in the text. Moreover, in order to be practical and to be capable of performing with some hope of success, LE systems must be furnished with a large-size lexicon, covering a realistic vocabulary, and providing the types of linguistic knowledge required for the applica tion. A survey of the types of linguistic knowledge needed for different systems is found in the final report of the EC Eurotra-7 project (Heid, McNaught, 1991). If, in addition, we do not want to build a new lexicon for each new application or system, we need to build large, generic and reusable lexicons (Calzolari, 1991), from which the required data in the required format can be extracted by different applications through appropriate filtering and conversion procedures. It is a matter of fact that, given the present state-of-the-art, the linguistic information required for real-life applications, which has to be encoded in a Computational Lexicon, possibly in a standardised or normalised way, can be very complex and difficult to acquire/gather, to structure, and to represent in a formal way. Even though many steps forward have been made in the last ten years as regards computational lexicons, we are still in a position to reiterate that the lexicon is a major bottleneck for natural language processing (NLP) systems. The capa bilities of NLP systems were and have remained weak because of the labour intensive nature of encoding lexical entries. After approximately ten years of acquiring (semi-)automatically lexical/linguistic information from machine-readable dictionaries (see Calzolari, Briscoe, 1995 for an overview of the ACQUILEX project which was started exactly from this hypothesis of work), we can clearly see not only the strong points but also the intrinsic limitations of this