Focused Mining of University Course Descriptions from Highly Variable Sources (Abstract Only)
Thomas Effland · 2015
Finding topically relevant content from disparate sources on the Web requires robust techniques due to the variability of sites. A focused web crawler is a type of crawler that attempts to make predictions about page relevance and traverse the web efficiently. In this work, we attempt to design a novel system of focused crawling tailored to identifying and extracting semantically similar topical information from disparate but known seed domains with highly variable structure that do not reference each other. We first extract rich predictive features from web pages. We then utilize Weakly-Supervised Machine Learning techniques to predict the link distance of current pages to target pages by employing two separate Random Forest classifiers that rank the current page and potential relevance gain of hyper-links. We use these page representations and rankings to efficiently tunnel through irrelevant pages and reach target pages with more optimal path traversals.