A similarity-based agent for internet searching
Tony G. Rose, Peter Wyard · 1997
This paper describes the development of a new Internet Information Agent (IIA) that uses similarity-based methods to search the Internet. The Agent works by analysing a sample of the type of text that is known to be of interest to the user. It then extracts a number of linguistic features and stores these as a feature vector that is used to describe the content of the document. This data is then used as input to a range of similarity metrics that allow the agent to compare new texts with the original and thereby acquire "more of the same".