UCEJ Database Refinement and Applicability Proof
Antoine Bossard, Keiichi Kaneko · 2019
The representation by computer systems of Chinese characters is an ongoing issue: it is still impossible to use some of them on a computer. Several encoding solutions have been proposed over the years, with most notably two approaches: the unifying approach followed by Unicode which aims at covering all the glyphs known to mankind, and the non-unifying approach followed by "local" encodings such as Shift-JIS and EUC-JP in the case of Japanese. In previous works, we have proposed an unrestricted character encoding for Japanese (UCEJ) so as to address the issues faced by the other encodings. In this paper, we propose a refinement to the realisation method of the character database on which UCEJ is based, and then show the applicability of the UCEJ encoding. To this end, we first describe a proof of concept UCEJ application (viewer and converter) and finally quantitatively compare UCEJ against Unicode with respect to memory size requirements. Not only does UCEJ features essential improvements over Unicode regarding the code structure and features overall, but we show that the induced memory size overhead can be almost eliminated under certain conditions.