Indexing and querying linguistic metadata and document content
Niraj Aswani, Valentin Tablan, Kalina Bontcheva, Hamish Cunningham · Amsterdam studies in the theory and history of linguistic science. Series 4, Current issues in linguistic theory · 2007
The need for efficient corpus indexing and querying arises frequently both in machine learning-based and human-engineered natural language processing systems. This paper presents the ANNIC system, which can index documents not only by content, but also by their linguististic annotations and features. It also enables users to formulate versatile queries mixing keywords and linguistic information. The result consists of the matching texts in the corpus, displayed within the context of linguistic annotations (not just text, as is customary for KWIC systems). The data is displayed in a graphical user interface, which facilitates its exploration and the discovery of new patterns, which can in turn be tested by launching new ANNIC queries. 1