Variability as a measure of semantic structure for document storage and retrieval
Alberto J. Cañas · 1985
The introduction of office support technologies has created a situation in which unstructured documents are increasingly likely to be stored and retrieved from computers. Unstructured documents require different organizations and tools than those employed in existing structured databases (e.g. relational). This thesis is based on the notion that a better understanding of the way people classify and search for information may be instrumental in designing more effective systems. A field study carried out in four offices to investigate current handling of unstructured documents is described. A conceptual model of the process of document storage and retrieval is then proposed. Two experiments were conducted to test the model and to compare systems in manual and automated environments. The results of the experiment on the manual system were supportive of the basic approach of the model and the use of variability as a measure of semantic structure. The manual experiment also provided a retrieval performance baseline to evaluate a prototype computer-based structure editor in the second experiment. The results of this second experiment indicate that measures of semantic structure are more informative than simple measures of physical structure.