An Interactive Approach to Development of English-Tamil Machine Translation System on the Web.
Vasu Renganathan · 1995
This paper illustrates a research being pursued on development of English-Tamil machine translation system. This is a rule-based system containing around five thousand words in lexicon, and a wide range of transfer rules written in Prolog encompassing basic English structures mapped to corresponding Tamil structures. Both rule base and the lexicon of this system are built in such a way that the users can update the scope of this system interactively by adding words into lexicon and rules into rule-base. Translating both colloquial and technical English into Tamil with a computer essentially involves construction of the two basic blocks namely the lexicon and rule base. Construction of online lexicon requires codification of grammatical information in two different ways. One by coding a minimal-set of information about grammatical categories of head and target words and the other by including an extensive information involving semantic and syntactic properties of words. The former type of lexicon is sufficient for translating technical, colloquial and news documents, where as the latter type of lexicon is mandatory for translating literary texts comprising fiction, poems, biographies etc. The system demonstrated here is built with the former type of lexicon containing a minimal set of grammatical information about head words of both English and Tamil. The other significant component of any machine translation system is building rule base that maps the structures of both source and target language. Any ideal system should be capable of accommodating not only the basic structures of source language, but also a wide variety of complex structures accounting for all kinds of ambiguous interpretations. The