A Software Toolkit for Sharing and Accessing Corpora Over the Internet

Saturnino F. Luz · 2000

This paper describes the Translational English Corpus (TEC) and the software tools developed in order to enable the use of the corpus remotely, over the internet. The model underlying these tools is based on an extensible client-server architecture implemented in Java. We discuss the data and processing constraints which motivated the TEC architecture design and its impact on the efficiency and scalability of the system. We also suggest that the kind of distributed processing model adopted in TEC could play a role in fostering the availability of corpus linguistic resources to the research community. 1. Background Corpus linguistics has gained increasing importance in both theoretical and computational linguistics. Witness the considerable volume of corpora and software tools for cor-pus processing currently on offer from linguistic data asso-ciations in Europe and the US. Very large corpora such as the Bank of English have been used for some time in the area of lexicography for analysing collocation patterns (Clear, 1993), normally in

Read the paper · More papers on PaperTik