On OLAP Data Model Driven Approach to Process Statistical Tables
Wo-Shun Luk, Patrick Leung · 2006
Statistical tables belong to an important subset of tables published in the Web, because they represent up-to-date, vital information sources for decision makers. These tables are often carefully designed for easy reading by analysts, and then mechanically produced by an OLAP database system. The general practice of extracting attribute-value pairs from statistical tables does not ensure high accuracy when they are used as a database for an information retrieval system. In this paper, we show how a human may visualize a statistical table as an multidimensional object, defined by a suitably modified OLAP model. In this way, the keywords are classified into semantically distinct groups, i.e., dimension hierarchies, without any ontological knowledge or resorting to machine learning. A prototype system which mimics the human reasoning for table processing has been implemented. Experiments on 150 randomly chosen tables from statistics Canada have confirmed the validity of this approach.