Stratified k-means clustering over a deep web data source

Tantan Liu, Gagan Agrawal · 2012

This paper focuses on the problem of clustering data from a {\em hidden} or a deep web data source. A key characteristic of deep web data sources is that data can only be accessed through the limited query interface they support. Because the underlying data set cannot be accessed directly, data mining must be performed based on sampling of the datasets. The samples, in turn, can only be obtained by querying the deep web databases with specific inputs.

Read the paper · More papers on PaperTik