Evaluation of Different Similarity Measures for the Extraction of Multiword Units in a Reinforcement Learning Environment
Gaël Dias, Sérgio Nunes · 2004
In this paper, we present an application of Genetic Algorithms to extract Multiword Units (i.e.complex lexical units such as compound nouns, idiomatic expressions or phrase templates).For that purpose, a fitness function will be defined whose maximization will serve as a basis for the identification of pertinent word -grams (i.e ordered vectors of words) based on different similarity measures.Finally, we will provide an experiment realized over an English Linux Manual that evidences promising results.