The Contextual Analysis of Chinese Sentences with Punctuation Marks

Hsin‐Hsi Chen · Literary and Linguistic Computing · 1994

By corpus analysis, about 75% of Chinese sentences are composed of more than two sentence segments separated by commas or semicolons. A segment may be a sentence, a noun phrase, a verb phrase, an adjective phrase, an adverbial phrase, or a prepositional phrase. An NP segment may service as a subject of the next segment or an object of the previous segment. The empty category pro may also appear in the VP segment. The maximal freedom of the uses of pros, the large number of segments, the various segment types, and the associativity problem make sentence parsing difficult. Few parsing systems deal with these problems. This paper regards a segment as a basic parsing unit. And it uses punctuation marks, categories of segments, linking elements, topic chains and some heuristic rules to link the segments into meaningful units. The pro resolution and the segment linking are useful for practical applications. Machine translation is a typical example.

Read the paper · More papers on PaperTik