A Study on the TEI Standard Annotation for Tibetan Corpus
Dingguo Gao · Zhongwen xinxi xuebao · 2011
Large-scale real text processing has become a hotspot in the language information processing.To annotate the Tibetan Corpus is very important for the research on Chinese-Tibetan machine translation,information retrieval,text data mining and dictionary compilation.To facilitate the data exchange and sharing,this paper studies on on adopting the TEI coding for Tibetan corpusannotation,including the text attribute information and structure information.