Structure and interoperability of online encyclopedic content in the field of technology
Ivan Smolčić · 2020
The implementation of encyclopedic works as contemporary network projects enabled the upgrade of these information systems in an epistemological sense, ie the necessary redefinition of lexicographical and encyclopedic activities. While retaining the basic (traditional) determinants of the encyclopedic concept, such as accuracy, objectivity and relevance, online encyclopedic editions go beyond the traditional ones in their content creation, removed scope limits (thanks to digital content), enhanced content retrieval and user adaptability. This work addresses in detail the issues of structure and interoperability of encyclopedic content, which enable the encyclopedic content to interact and function in the global information system. An overview of the main features and the application of the structure and interoperability of encyclopedic content in the network space is provided. In order to draw conclusions about the possibilities of structuring encyclopedic content on the basis of its informative nature, a quantitative and qualitative content analysis was conducted in order to record unified factual data suitable for structure creation. A total of 455 encyclopedic articles from 10 separate online encyclopedic projects in Croatian, English and German were examined. Entities recorded in the content analysis of encyclopedic texts were divided into three categories of encyclopedic articles: persons, organizations and general terms. The article category for persons contains a total of 33 types of entities suitable for structuring, the article category for organization marks 23 entity types, and the article category for general terms contains only seven entity types. In order to facilitate automatic identification of named entities in encyclopedic texts, the effectiveness of Named Entity Recognition tools on encyclopedic texts in the technical field has been tested. Three tools (applications) were used for testing: CroNER and ReLDI for the Croatian samples and the Stanford CoreNLP tool for the samples in English. Based on the parameters derived from the tool-marked text, using a mixed sample of articles from multiple encyclopedic publications, standard performance measures (precision, response, accuracy, F measure) were calculated for each tool, providing evaluation of the tools for individual categories and types of names. On the basis of the content analysis, the method of structuring the encyclopedic content and the use of that structure in achieving interoperability by developing a special model are presented. The model represents a central database that binds the structure of multiple editions and enables it to be exchanged to all related online encyclopedic editions by mapping the elements of the structure. Interoperability has been proven via experimental method by testing models and detecting changes. Ultimately, this PhD thesis draws a conclusion regarding all its constituents, that is, points to encyclopedic virtues through the settings of the encyclopedic concept, the role of structure and the possibility of interoperability of online encyclopedic content.