Sophisticated Information Gathering in a Marketplace o f Information Providers
Victor Lesser, Bryan Horling, Anita Raja, Thomas A. Wagner · 2004
BIG is a sophisticated, web-based, information gathering agen t that recommends software packages. BIG plans, locates a n d processes free-format WWW documents via natural language processing and other text extraction techniques. BIG uses t h e processed information to create models of software products a n d then compares the models to the client’s criteria a n d recommends a product for the client to purchase. BIG also uses the information extracted during the search to adapt and ref ine the search process itself – for example, deciding to gather m o r e information on a particularly highly referred product. This paper discusses techniques used by BIG to control: 1) the a m o u n t of monies spent in acquiring information from sites that charge a fee for accessing their information, 2) the balance between t h e scope/coverage of information gathered and the precision o f resulting decision, and 3) the end-to-end time that t h e information gathering and processing activities will take. As p a r t of this discussion, we present the DTC (Design-to-Criteria) scheduler and the domain-independent activity representat ion, called TAEMS, that is used to describe and quantify BIG’s p rob lem solving activities for the scheduler. We also present experimental results showing how BIG reconfigures its activities to meet resource and cost constraints. 1 This material is based upon work supported by the Department of Commerce, the Library of Congress, and the National Science Foundation under Grant No. EEC-9209623, and the National Science Foundation under Grant No. IRI-9523419, and the Department of the Navy and Office of the Chief of Naval Research, under Grant No. N00014-95-1-1198. The content of the information does not necessarily reflect the position or the policy of the Government or the National Science Foundation and no official endorsement should be inferred.