Document Search Images in Text Collections for Restricted Domains on Websites
Pavel Makagonov, Celia B.Reyes E., Grigori Sidorov · IGI Global eBooks · 2011
The main idea of the authors’ research is to perform quantitative analysis of a text collection during the process of its preparation and transformation into a digital library for a website. They use as a case study the digital library of the website on Mixtec culture that we maintain. The authors propose using the concept of the text document search image (TDSI). For creating TDSIs they make analysis of word frequencies in the documents and distinguish between the Zipf’s distribution that is typical for meaningful words and distributions approximated by an ellipse typical for auxiliary words. The authors also describe some analogies of these distributions in architecture and in urban planning. We describe a toolkit DDL that allows for TDSI creation and show its application for the mentioned website and for the corpus of dialogs with railway office information system.