AUTOMATIC ACQUISITION OF BILINGUAL LANGUAGE RESOURCES

Nikos Mastropavlos, Athena Ric, Vassilis Papavassiliou · 2012

This paper discusses methods for automatic acquisition of bilingual corpora from the Web. Given the vast number of documents available online, the Web could be considered an excellent pool for extraction of valuable data for linguistic purposes. Therefore, methods for creating such corpora, especially when targeting less-resourced languages like Greek, can be of great value. Besides presenting a general workflow for constructing collections from the Web, this article describes our work to produce collections of English/Greek comparable documents in the “Political News”, “Technological News”, “Sport News”, and “Renewable Energy” domains and parallel resources in the “Environment” and “Labour Legislation” domains.

Read the paper · More papers on PaperTik