BAD LUC$@$WMT 2016: a Bilingual Document Alignment Platform Based on Lucene
Laurent Jakubina, Phillippe Langlais · 2016
We participated in the Bilingual Document Alignment shared task of WMT 2016 with the intent of testing plain cross-lingual information retrieval platform built on top of the Apache Lucene framework.We devised a number of interesting variants, including one that only considers the URLs of the pages, and that offers -without any heuristic -surprisingly high performances.We finally submitted the output of a system that combines two informations (text and url) from documents and a post-treatment for an accuracy that reaches 92% on the development dataset distributed for the shared task.