Methods and Development of Chinese Word Tokenization
Zhenghan Fang · Applied and Computational Engineering · 2024
Chinese word tokenization is an important task in natural language processing and has undergone significant evolution with artificial intelligence. This paper reviews the historical progression and contemporary methodologies for Chinese word segmentation. The paper examines the traditional character-based approaches, which rely on dictionaries and pattern matching, and transition into machine learning-based techniques that utilize statistical models and neural networks. A particular focus is given to the recent developments in deep learning, including the application of recurrent neural networks (RNNs), long short-term memory networks (LSTMs), and transformer models like BERT. The review also highlights innovative approaches such as memory networks and sub-character tokenization, which have shown promising results in improving segmentation accuracy and computational efficiency. Furthermore, the paper discusses the challenges faced in tokenization, such as handling out-of-vocabulary words and the integration of syntactic and semantic information. The paper concludes with insights on the future directions of Chinese word tokenization, emphasizing the potential of unsupervised learning and the need for more robust evaluation frameworks.