Statistic Chinese New Word Recognition by Combing Supervised and Unsupervised Learning
Fei Wang · 2019
This paper proposed a method using statistic methods with untagged and tagged data to detect Chinese new words. Chinese sentences have no blank to segment the characters to build words. So Chinese word segmentation has been a hot research field for a long time. Chinese word segmentation needs tagged data and dictionary. The dictionary should always be updated, thus ththereere will be less out of vocabulary words in the model. It will cost a lot of time to collect new words manually, and it has many limitations such as the mount of the words. Term frequency, pointwise mutual information and entropy are applied to detect new Chinese new words. They are based on the untagged data by statistic methods. Hidden markov model is used to segment Chinese sentences and detect new words. Hidden markov model need tagged data to train the model. Hidden markov model can learn the probability of a character's position of a word such as begin, middle, end or single. We combine these two methods: one method to use the untagged data and an other to use the tagged data. The untagged data has many resources and it does't cost any human operation, but it is not so accurate. The tagged data is tagged by human, so it is accurate and reliable, but it takes a lot time to tag it, only a few of data is free to use to train the method. In short, we try to use tagged data and untagged data at the same time by combing pointwise mutual information, entropy and hidden markov model. The experiment shows the new word candidates number is reduced from 306% to 173% compared to the training dictionary number, but the real words number is only reduced 4%.