Research on Tibetan News Sites’ Web Crawler and Search Engine
Zhiqiang Han, Guixian Xu, Wei Sun · Advances in intelligent systems research/Advances in Intelligent Systems Research · 2015
In this paper, researchers detailedly introduce the features of Tibetan language and related technologies that researchers use to deal with Tibetan news web pages with computers.To get the content of the Tibetan news, researchers used web crawler to download Tibetan news pages which were the bases of this project.Researchers used an open source web crawler named scrapy and rewrote the crawl part to make the crawler work more accurately and efficiently.To search the Tibetan content in a way, researchers defined and counted every statistical data that was useful and helpful to enhance the performance of the search engine.Researchers used solr, another open source software, as the user interface of this system.The crawler and search engine were combined by the web pages to provide the data retrieval service.Comparing with other works, researchers' work adopted a safe and stable enough framework to enhance the user experience in using Tibetan search engine.Researchers' work played a positive role in the spread of Tibetan culture and promoted the development of the Tibetan language news in the field of search engines.