Research on Chinese Complex Grammar Recognition Utilizing K-means Clustering and BCCNN Algorithm

Di Wu · 2024

In the modern field of natural language processing, accurate identification of complex grammar structures is crucial for achieving efficient language understanding. This paper proposes a new method that combines K-means clustering and Boundary-Concatenation Convolutional Neural Network (BCCNN) to improve the accuracy of recognizing complex Chinese grammar. Initially, the K-means algorithm is employed to preprocess sentences in the corpus, clustering them based on grammar complexity. This step effectively aggregates structurally similar sentences, providing a finer data foundation for subsequent deep learning model training. Subsequently, a BCCNN based on boundary concatenation strategy is designed to enhance the model's sensitivity to sentence boundaries and grammar structures through specific merging mechanisms. Experimental results demonstrate that compared to traditional CNN models and other baseline models, this method exhibits higher accuracy and robustness in recognizing complex Chinese grammar structures. The experiments reveal that this method not only effectively identifies and parses complex Chinese grammar structures but also significantly improves the ability to process complex sentences, offering an effective technical approach for a deeper understanding of the complexity of Chinese grammar.

Read the paper · More papers on PaperTik