Natural Language-Enabled Data Modeling
Alexander Hars · Journal of Database Management · 1998
Although data modeling has been an area of intensive research, there is a lack of operational procedures for measuring model quality and for integrating large-scale models. Progress has been limited because the meaning associated with the individual elements of a model needs to be taken into account. This meaning rarely is made explicit and therefore not directly available for validation and integration procedures. In this paper we show that natural language processing techniques based on a custom-built dictionary can be used to interpret data models and leverage meaning for validation and integration procedures. We describe a general-purpose dictionary which contains syntactic and semantic word categories for 23.000 English words. We show how the syntactic and semantic information in the dictionary can be used to detect semantic inconsistencies, reject inconsistent naming, identify unregistered abbreviations and detect synonym candidates. To prove feasibility of our approach, we describe a prototype tool which has been implemented on a standard personal computer.Request access from your librarian to read this article's full text.