Computing Linear Discriminants for Idiomatic Sentence Detection
Jing Peng, Anna Feldman · 2009
In this paper, we describe the binary classication of sen- tences into idiomatic and non-idiomatic. Our idiom detection algorithm is based on linear discriminant analysis (LDA). To obtain a discriminant subspace, we train our model on a small number of randomly selected idiomatic and non-idiomatic sentences. We then project both the train- ing and the test data on the chosen subspace and use the three nearest neighbor (3NN) classier to obtain accuracy. The proposed approach is more general than the previous algorithms for idiom detection | neither does it rely on target idiom types, lexicons, or large manually annotated corpora, nor does it limit the search space by a particular linguistic con- struction.