Understanding, interpreting and querying Web statistical tables

Chi Hang Leung · Summit (Simon Fraser University) · 2005

Extraction of information from tables published on the Web is made less complicated because of easy identification of the text inside a table cell.In this thesis, we propose, and have implemented, a scheme which not only understands the contents in a statistical table, but is also able to convert them into a multidimensional database which can then be fed into an off-the-shelf system for querying and data integration.By carefully interpreting the intention of the table author via the visual cues embedded into the HTML text, and the layout design of multidimensional database modelling techniques, our system can successfully classify the keywords into semantically distinct dimension hierarchies, without any domain-specific knowledge, or machine learning.Experiments on a set of real-life statistical tables have confirmed the validity of this approach.Experiments on a set of real-life statistical tables have confirmed the validity of this approach.iii Dedication To my dearest parents

Read the paper · More papers on PaperTik