Non-word Strings in Open Source Word Vectors

Xingyuan Chen, Peng Jin, Jiuhua Zhang, Caiming Liu · 2022 IEEE 10th Joint International Information Technology and Artificial Intelligence Conference (ITAIC) · 2022

The Chinese open source word vector corpus is an important corpus for current research and application. Due to the imperfect word segmentation method, it contains a large number of non-word strings, which brings difficulties to the application. Investigating and analyzing such strings in open source word vectors can provide valuable reference information on how to handle them. This paper uses a widely used open source word vector corpus containing more than 8 million: the Tencent word vector corpus, and conducts a preliminary analysis on the non-word situation by combining manual annotation and rule-based query technology. Experiments have found that the open this source word vector corpus contains a large number of non-word strings, especially the strings in the word vector entries with three-character words and above, about 80% are not words. Obviously, this brings a lot to the use of Huge damage. We analyze the specific situation of this kind of harm with the example of word similarity application.

Read the paper · More papers on PaperTik