Termediator II
Owen Riley, Jessica Richards, Joseph J. Ekstrom, Kevin Tew · 2014
We report on Termediator II, an application designed to identify potentially confusing terms. Termediator I focused on identifying synonymous terms whereas this work, Termediator II, focuses on identifying polysemous terms. Using an expanded collection of 399 glossaries, we combine hierarchical clustering algorithms and text similarity measures to assign each terms a numeric value indicating its degree of polysemy. Cosine, latent semantic indexing (LSI), and latent Dirichlet allocation (LDA) text similarity measures are evaluated using hierarchical agglomerative clustering with complete and average linkage types. To improve results, we combined bodies of knowledge (BOKs) with the glossaries to create an enhanced training corpus for LSI and LDA. We introduce the convergence value as a new generic metric of polysemy. Polysemous terms are identified by sorting the glossaries by cluster quantity at the convergence value. The similarity measure and linkage type combinations produced slightly different but effective lists of highly polysemous terms.