Automatic extraction and alignment of multiword expressions from English-Chinese comparable corpus
Yuan Qi · Computer Engineering and Applications Journal · 2010
Multiword Expressions(MWE) are important for practical applications,such as machine translation(henceforth,MT),multilingual information retrieval,data mining and other natural language processing.A method of combining semantic template and statistical tool is proposed for automatically extracting native English MWE from three-tuple comparable corpus.The saurus-based and distributional methods are harnessed to calculate the semantic relations between words for improving MWE coverage.GIZA++ is executed to align words at sentence level,aiming at obtaining Chinese MWE candidates.For each native English MWE,all of the Chinese MWE candidates are collected and sorted according to their co-occurrence affinity.Only the top one is accepted as true Chinese translation of the given English MWE.Experimental results show the proposed technique improves MWE extraction and alignment efficiently.