Tagging Email with 8 Tagsets: lessons on evaluation
Eric Atwell, George Demetriou, Hughes, John, Schiffrin, Mandy, DC Souter, Sean Wilcock · White Rose Research Online (University of Leeds, The University of Sheffield, University of York) · 1997
An email server has been established at Leeds University to assist grammatical analyses of English texts. Text can be emailed to the server [email protected] at the Leeds Centre for Computer Analysis of Language and Speech (CCALAS). It will be returned with Part-of-Speech (POS) tags incorporated as required. For further details, send an email with subject 'help' and a blank message-body to receive the instruction file; or see WWW URL http://agora.leeds.ac.uk/amalgam/ The automatic POS-tagger was developed in the Automatic Mapping Among Lexico-Grammatical Annotation Models (AMALGAM) project. There are several other taggers; but our service is different because it is available by email and offers a choice from eight standard POS-tag categories. Project AMALGAM aimed to create a set of mapping algorithms between the grammatical annotation schemes used in the main ICAME corpora, to allow English Corpus researchers to compare and use each other's corpus annotation schemes. In fact, we found that a more practical and effective approach was to build a multitagger, capable of annotating text with tags from any (or all) of the eight English Corpus tagsets: 1.Brown Corpus (Brown) 2.International Corpus of English (ICE) 3.Lundon-Lund Corpus (LLC) 4.Lancaster-Oslo/Bergen Corpus (LOB) 5.UNIX parts (Parts) 6.Polytechnic of Wales Corpus (POW) 7.Spoken English Corpus (SEC) 8.University of Pennsylvania Corpus (UPenn) We have compiled a detailed description of each of these tagsets on our WWW pages, including example words for each tag, and an example text tagged according to each of the schemes. We have also compiled parsed versions of the example text, again parsed according to a range of different parsing schemes. This is a unique resource for the comparison and evaluation of underlying grammatical classification schemes. The current service is an experimental prototype and is available free of charge. We are also investigating an interactive form-based version directly usable from our WWW pages; but we think that the email service may be better for many users. We do NOT keep a copy of your text; this is to avoid possible copyright problems, and also to avoid clogging up our diskspace! We DO keep a record of the size and source-address of email; we have followed up some users with a short questionnaire asking their opinions, and some background on what they have used the system for. The service was launched in January 1997, and we have been both surprised and gratified by the volume of use and the wide range of applications, including reviews in Machine Translation Review and IMPACT Engineering and Physical Sciences Research Council newsletter. We recommend that other ICAME researchers `show off their wares' to the world over the Internet, to allow a wide range of potential users to evaluate them. We end with some warnings about possible pitfalls: we learnt the hard way that, if amalgam-tagger is open to use by absolutely anyone over the internet, it needs to be fool-proof! Rather than take up your valuable time during the workshop, We urge participants to try out our `online demo' beforehand. If you try out [email protected] and return the follow-up questionnaire before May 16th, we can include your opinions in our Evaluation! References: see http://agora.leeds.ac.uk/amalgam/