Solving CAPTCHAs Automatically for Web Crawling

Swapnil S.Mane, Mayuri Lokare · International journal of advance research and innovative ideas in education · 2017

Web crawler is an automated script which browses the World Wide Web in methodical, automated manner. This process is called as web crawling or web spidering. Web crawler is also known as web spider or web robot. Crawler needs to catch web pages frequently for updating the data. But by this process, performance and speed may get affected. Crawler can not retrieve data in a great depth. Hence, to reduce load and also for authentication, web server requests web crawler to verify or cross check themselves against CAPTCHAs. Today, more than two billion web pages are there on web server. One human can not enter CAPTCHAs for every time while entering to required web page. Hence, to solve CAPTCHAs automatically we explained a system for text or image recognition from CAPTCHA images. Our focus is on reliable and firm text or image extraction, recognition, putting resolved CAPTCHA to crawler system to continue the web crawling process without involving human directly.

Read the paper · More papers on PaperTik