Corpus-Based Adaptation Mechanisms for Chinese Homophone Disambiguation.
Chao-Huang Chang · 1993
Based on the concept of &dwectwnal converswn and automallc evalua&on, we propose two useradaptalcon macbantams, character-preference learnin and pseudo-word learmnf, for resolwn Chinese homophone ambfuihes m syllable-to-character con- version. The 191 Urnted Daily corpus of approzimatel 10 milhon Ch,nese characters s used for traction of 10 reporter-specific article databases and for computatwn of word frequencies and character bi- :rarns. Ezperarnenla show that 0.$ parcertl (testm sets] to 71.8 laercent (tramm sets) of conversmn rots ca be ehrnmated throuh the proposed rnechamama.