A Study on the TEI Standard Annotation for Tibetan Corpus

Dingguo Gao · Zhongwen xinxi xuebao · 2011

Large-scale real text processing has become a hotspot in the language information processing.To annotate the Tibetan Corpus is very important for the research on Chinese-Tibetan machine translation,information retrieval,text data mining and dictionary compilation.To facilitate the data exchange and sharing,this paper studies on on adopting the TEI coding for Tibetan corpusannotation,including the text attribute information and structure information.

Read the paper · More papers on PaperTik