ODIN: A Model for Adapting and Enriching Legacy Infrastructure
William Dodge Lewis · 2006
The Online Database of Interlinear Text (ODIN)1 is a database of interlinear text “snippets”, harvested mostly from scholarly documents posted to the Web. Although large amounts of language data are posted to the Web as part of scholarly discourse, making the existing “e-Linguistic in-frastructure ” surprisingly rich, most linguistic data avail-able on the Web exists in legacy formats, is highly display-centric, and is often difficult to locate or interoperate over. ODIN seeks to leverage this existing infrastructure into a rich, searchable, and interoperable resource by converting readily available semi-structured data to content-centric, searchable formats. To do this, ODIN mines scholarly pa-pers and webpages for instances of linguistic data, focusing mostly on interlinear texts, extracts them, identifies source languages, and makes the instances available to search. Through ODIN’s standard search feature, users can locate data by language name or Ethnologue code, and display lists of data by document for languages of interest. The newer Advanced Search feature allows users to locate in-stances by grammatical markup that is used (e.g., NOM, ACC, ERG, PST, 3SG), and by linguistic constructions (e.g., passives, conditionals, possessives, raising constructions, etc.). The latter are made possible through additional en-richment of discovered data using automated statistical tag-gers and parsers. 1 The ODIN Vision The Online Database of Interlinear Text (ODIN) is a database of interlinear text “snippets”, harvested mostly from scholarly documents posted to the Web. ODIN was developed as part of the greater effort within the GOLD