The WATCHER Project: Building an Agent for Automatic Extraction of Language Resources from the Internet

Kyriakos Sgarbas · Literary and Linguistic Computing · 2003

The WATCHER project aims to automate the extraction of language resources from the Internet via an intelligent agent called the ‘WATCHER’. This agent (in its final form) will be able to actively search and collect subject-specific and language-specific texts and build corpora and lexicons from them. Although the resources will still have to be checked for validity after their collection, the proposed method requires the minimum of human interaction. Apart from its ability to collect these resources automatically, the WATCHER will also be able to track the evolution of a target language over time by collecting resources annually and presenting their analysis in annual reports. The WATCHER is still under development. This paper presents an overview of its architecture and functionality, and reports recent progress.

Read the paper · More papers on PaperTik