An arabic lexical database to support natural language processing
Awni Hammouri · 1995
We are building a lexical database for Arabic to support information retrieval, parsing, and text generation. It contains information about words and phrases in the computer sublanguage. For each entry there is a part of speech and morphological information (information about roots and stems). For verbs we give information about the number and type of arguments. For nouns we list gender (masculine vs. feminine), and number (singular, dual, plural). We also indicate whether each noun is abstract or concrete, human or animate or inanimate, and countable or mass. Adjective entries in the lexical database include information about morphology (comparative and masculine vs. feminine forms) and semantic categories (dynamic vs. stative, gradable vs. nongradable, inherent vs. noninherent). We also store information about lexical-semantic relations in the Arabic language. Nowadays the rapid development in computer technology makes computer applications and software systems available and easy to use in almost every field in the Arabic countries. But users of computers in the Arabic world need to know English because Arabic interfaces and Arabic natural language processing systems are not widely available. In order to develop these systems a computer lexicon of Arabic is badly needed. The Arabic lexicon as a research field has been studied in computational linguistics only in the last few years and thus research is still far from achieving these goals. This work also has significance for research in the Arabic language, and especially in Arabic lexicography. I want to study Arabic language problems and try to find possible solutions. I want to encourage computer scientists and linguists to work together on a universal lexicon for Arabic.