Almost-unsupervised cross-language opinion analysis at NTCIR-7

Taras Zagibalov, John Alexander Carroll · Figshare · 2008

Abstract We describe the Sussex NLCL System entered in the NTCIR-7 Multilingual Opinion Analysis Task (MOAT). Our main focus is on the problem of portability of natural language processing systems across languages. Our system was the only one entered for all four of the MOAT languages, Japanese, English, and Simplified and Traditional Chinese. The system uses an almost-unsupervised approach applied to two of the sub-tasks: opinionated sentence detection and topic relevance detection. Keywords: NTCIR, opinion analysis, relevance detection, cross-language portability. 1 Introduction In our entry to the NTCIR-7 Multilingual Opinion Analysis Task (MOAT) [1] we focused on developing a system that can be easily ported from one language to another.The most obvious way to design such a system is to make it as unsupervised as possible. Unsupervised systems derive all their information from raw, unannotated language data. Unsupervised techniques are promising for systems that need to be ported across domains, text types and languages. However, one of the main problems of unsupervised systems is that they are usually less accurate than supervised systems, which have access to annotated training data. Another problem is that unsupervised methods often need a large amount of raw data to be able derive any useful information. As well as not using annotated data, unsupervised systems typically also contain few built-in assumptions about how a language is structured, encoded for example as rules or via a lexicon. Rule-based systems might consist of hundreds of rules and writing these manually may be as costly in time and linguistic expertise as annotating a corpus to be used by a supervised system. We therefore use no language-specific rules or other kinds of processing which would render our system less portable.In this paper we present a system for opinion analysis which uses an unsupervised approach combined with a very limited amount of manual intervention. We participated in two of the MOAT sub-tasks: opinionated sentence detection (required for all participants) and topic relevance detection (an optional sub-task). We applied our system to all four of the languages, Japanese, English, and Simplified and Traditional Chinese.

Read the paper · More papers on PaperTik