The Model of Semantic Similarity Estimation for the Problems of Big Data Search and Structuring
Victoria V. Bova, Vladimir V. Kureichik, Dmitry Leshchanov · 2017
The main problem in the field of Big Data search and processing involves constantly growing complexity of its identification and structuring for the purpose of representation in the form suitable for understanding and further use. To solve this problem authors propose to use method of multilevel semantic net building to define connections between data meta-descriptions in large distributed information arrays. The semantic model developed on the basis of the method provides visibility and compact presentation of structure of semantic relations between mass data arrays elements. Semantic meta-descriptions are considered as sets of triples “subject-predicate-object” in terms of subject area ontology of distributed operative databases and the query. Authors propose the model to search and estimate semantically similar elements of distributed databases based on clustering of semantic nets represented as graph models on corresponding levels: subject area level, search profile level and document meta-descriptions level. The relevance (semantic similarity) estimation method is based on closeness assessment of data in distributed information arrays of document and query semantic nets. To analyze the developed method authors carried out a set of computational experiments. Obtained data proved theoretical significance and application perspective of such approach.