Multilingual Distributional Lexical Similarity
Kirk Baker · OhioLink ETD Center (Ohio Library and Information Network) · 2008
One of the most fundamental problems in natural language processing involves words that are not in the dictionary, or unknown words.The supply of unknown words is virtually unlimited -proper names, technical jargon, foreign borrowings, newly created words, etc. -meaning that lexical resources like dictionaries and thesauri inevitably miss important vocabulary items.However, manually creating and maintaining broad coverage dictionaries and ontologies for natural language processing is expensive and difficult.Instead, it is desirable to learn them from distributional lexical information such as can be obtained relatively easily from unlabeled or sparsely labeled text corpora.Rule-based approaches to acquiring or augmenting repositories of lexical information typically offer a high precision, low recall methodology that fails to generalize to new domains or scale to very large data sets.Classification-based approaches to organizing lexical material have more promising scaling properties, but require an amount of labeled training data that is usually not available on the necessary scale.This dissertation addresses the problem of learning an accurate and scalable lexical classifier in the absence of large amounts of hand-labeled training data.One approach to this problem involves using a rule-based system to generate large amounts of data that serve as training examples for a secondary lexical classifier.The viability of this approach is demonstrated for the task of automatically identifying English loanwords in Korean.A set of rules describing changes English words undergo when 4.1 Frequent English loanwords in the Korean Newswire corpus . . . . . . .5.1 Correlation between number of verb senses across five classification schemes120 5.2 Correlation between number of neighbors assigned to verbs by five classification schemes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .5.3 Correlation between neighbor assignments for intersection of verbs in five verb schemes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .5.4 An example contingency table used for computing the log-likelihood ratio 6.1 Number of verbs included in the experiments for each verb scheme . . . .6.2 Average number of neighbors per verb for each of the five verb schemes .6.3 Chance of randomly picking two verbs that are neighbors for each of the five verb schemes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .6.4 Examples of Subject-Type relation features . . . . . . . . . . . . . . . .6.5 Examples of Object-Type relation features . . . . . . . . . . . . . . . . .6.6 Examples of Complement-Type relation features . . . . . . . . . . . . . .6.7 Examples of Adjunct-Type relation features . . . . . . . . . . . . . . . .6.8 Example of grammatical relations generated by Clark and Curran (2007)'s CCG parser . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .6.9 Average maximum precision for set theoretic measures and the 50k most frequent features of each feature type . . . . . . . . . . . . . . . . . . . .6.10 Average maximum precision for geometric measures using the 50k most frequent features of each feature type . . . . . . . . . . . . . . . . . . . .6.11 Average maximum precision for information theoretic measures using the 50k most frequent features of each feature type . . . . . . . . . . . . . .6.12 Measures of precision and average number of neighbors yielding maximum precision across similarity measures . . . . . . . . . . . . . . . . . . . . .6.13 Nearest neighbor average maximum precision for feature weighting, using the 50k most frequent features of type labeled dependency triple . . . . .6.14 Average number of Roget synonyms per verb class . . . . . . . . . . . . .6.15 Nearest neighbor precision with cosine and inverse feature frequency . . .6.16 Coverage of each verb scheme with respect to the union of all of the verb schemes and the frequency of included versus excluded verbs . . . . . . .6.17 Expected classification accuracy.The numbers in parentheses indicate raw counts used to compute the baselines . . . . . . . . . . . . . . . . . . . .F.1 Average inverse rank score for set theoretic measures, using the 50k most frequent features of each feature type . . . . . . . . . . . . . . . . . . . .F.2 Average inverse rank score for geometric measures using the 50k most frequent features of each feature type . . . . . . . . . . . . . . . . . . . .F.3 Average inverse rank score for information theoretic measures using the 50k most frequent features of each feature type . . . . . . . . . . . . . .F.4 Inverse rank score results across similarity measures . . . . . . . . . . . .xiv F.5 Inverse rank score with cosine and inverse feature frequency . . . . . . .210 F.6 Nearest neighbor average inverse rank score for feature weighting, using the 50k most frequent features of type labeled dependency triple . . . . .