BUILDING AN ANNOTATED CORPUS AND A LEXICAL DATABASE OF MODERN HEBREW IN XML

Tsuguya Sasaki · Kyoto University Research Information Repository (Kyoto University) · 2004

In the present paper detailed schemes are proposed for an annotated corpus and a lexical database of Modern Hebrew. They are meant to be primary linguistic sources for more empirical studies of the grammatical and lexical structure of Modern Hebrew to shed light on aspects and phenomena hitherto unknown. XML (Extensible Markup Language) is chosen as their storage format because of its machine- and humanreadability, crossplatform-compatibility, crosslinguistic-compatibility, self-descriptiveness and capability of nesting structure. The corpus will be annotated in four levels, i.e., syntactically, morphosyntactically, lexically and morphologically. The lexical database will include modules of morphosyntax, inflection, word-formation and syntacticosemantics. The data will be recorded in Unicode, whether in Hebrew characters or in Latin transcription. Although countless numbers of revisions have been made since the idea of building these two sources in XML was first born a few years ago, what is proposed here is essentially by one individual. It might, therefore, need minor (or even major) revisions and/or expansions.

Read the paper · More papers on PaperTik