Extending database query models for intuitive data retrieval and analysis
Divyakant Agrawal, Ping Kun Tony Wu · 2007
With the ubiquitous adoption of the World Wide Web, an overwhelming amount of high quality, structured data becomes available on the Web. Unlike in the traditional context, Web databases are directly accessible by end users. To help ordinary users overcome the information overload in structured data, current databases need to improve the usability as well as the querying capability. This thesis makes an effort in this direction by introducing novel, user-friendly query types into existing data management systems and discussing techniques for enabling efficient and scalable implementations of these query operators. We focus on three query types, namely ranking queries, skyline queries and free-text (keyword) queries. Both ranking query and keyword query are the most intuitive query paradigms and have been the most popular way for ordinary users to express their information needs. However, due to the common coexistence of numerical and categorical attributes in the structured data, it is hard for an end user to specify the ranking function themselves. This is where the skyline query comes in. Skyline query and its variants return a set of interesting records from a database based on a set of attributes that a user wants to optimize. Any record in the skyline are guaranteed to be of interest assuming a monotonic preference. Our investigation covers two typical application scenarios: data retrieval and data analysis. We first study several distributed top-k algorithms, then we explore how scalable computation of skyline queries can be achieved via parallel processing as well as incrementally maintaining a computed result set. Finally, in the data analysis context, we show how skyline operations and keyword queries can be incorporated into OLAP systems to greatly enhance their query capability and usability.