A UNIVERSAL DATA MODEL FOR LINGUISTIC ANNOTATION TOOLS
Scott O. Farrar · 2006
An important goal of the E-MELD enterprise is to recommend bestpractice standards and resources in support of the digitization of endangered language data. As a result, there have been several proposals put forth at E-MELD sponsored events for best-practice data models, in particular, for dictionaries, paradigms, and interlinear text. While these proposals emphasize structural and encoding interoperability, by using XML and Unicode respectively, the resulting data models are not necessarily interoperable with respect to content. One suggestion for going beyond mere structural and encoding interoperability was to include in the various data models a reference to a common markup ontology, such as the General Ontology for Linguistic Description. As of yet, however, there have been few suggestions on how to implement such a model that emphasizes content interoperability via an ontology. This paper attempts to fill the gap by describing a common data exchange format to be useful for a variety of data digitization tools.