Automatic detection of English inclusions in mixed-lingual data with an application to parsing

Beatrice Alex · ERA · 2008

The influence of English continues to grow to the extent that its expressions have begun to permeate the original forms of other languages. It has become more acceptable, and in some cases fashionable, for people to combine English phrases with their native tongue. This language mixing phenomenon typically occurs initially in conversation and subsequently in written form. In fact, there is evidence to suggest that currently at least one third of the advertising slogans used in Germany contain English words. The expansion of the Internet, coupled with an increased availability of electronic documents in various languages, has resulted in greater attention being paid to multilingual and language independent applications. However, the automatic identification of foreign expressions, be they words or named entities, is beyond the capability of existing language identification techniques. This failure has inspired a recent growth in the development of new techniques capable of processing mixed-lingual text. This thesis presents an annotation-free classifier designed to identify English inclusions in other languages. The classifier consists of four sequential modules being pre-processing, lexical lookup, search engine classification and post-processing. These modules collectively identify English inclusions and are robust enough to work across different languages, as is demonstrated with German and French. However, its major advantage is its annotation-free characteristics. This means that it does not need any training, a step that normally requires an annotated corpus of examples. The English inclusion classifier presented in this thesis is the first of its type to be evaluated using real-world data. It has been shown to perform well on unseen data in both different languages and domains. Comparisons are drawn between this system and the two leading alternative classification techniques. This system compares favourably with the recently developed alternative technique of combined dictionary and n-gram based classification and is shown to have significant advantages over a trained machine learner. This thesis demonstrates why English inclusion classification is beneficial through a series of real-world examples from different fields. It quantifies in detail the difficulty that existing parsers have in dealing with English expressions occurring in foreign language text. This is underlined by a series of experiments using both a treebank-induced and a hand-crafted grammar based German parser. It will be shown that interfacing

Read the paper · More papers on PaperTik