Multilingual hyperdocument recognition: a document mining approach
Tuan Dang Nguyen, Khaldoun Zreik · 2004
This paper suggests a new distributive analysis approach to retrieve multilingual hyperdocument. To learn about the number of languages involved in a Web site, a set of general and computational knowledge is used, which is completely independent of the linguistic domains. The mining process considers three main stages: preprocessing (hyperdocument vectoring), processing (clustering), post processing (clusters pruning). A prototype of this system has been developed and tested two clustering approaches on several international Web - sites with very high and completely satisfactory performances.