Hidden Database Research and Analytics (HYDRA) System .
Yachao Lu, Saravanan Thirumuruganathan, Nan Zhang, Gautam Das · IEEE Data(base) Engineering Bulletin · 2015
A significant portion of data on the web is available on private or hidden databases that lie behind formlike query interfaces that allow users to browse these databases in a controlled manner. In this paper, we describe System HYDRA that enables fast sampling and data analytics over a hidden web database with a form-like web search interface. Broadly, it consists of three major components: (1) SAMPLE-GEN which produces samples according to a given sampling distribution (2) SAMPLE-EVAL that evaluates samples produced by SAMPLE-GEN and also generates estimations for a given aggregate query and (3) TIMBR that enables fast and easy construction of a wrapper that models both input and output interface of the web database thereby translating supported search queries to HTTP requests and retrieving top-k query answers from HTTP responses.