Design and Implementation of Web Information Extraction Platform Based on Rule Engine
Zhu Yan · Journal of Beijing City University · 2010
Information extraction is an important approach of data mining and knowledge discovery,accurate and valid Internet data extraction based upon rule engine as well as automation of the action are the key to knowledge discovery.This paper develops a general text information retrieval platform,using several kinds of information matching techniques to extract data from network data source and adopt processing rules to automatically and intelligently handle information.The platform is implemented using Eclipse RCP;features are implemented as Plug-ins and business logic is embodied as rules.The advantages of the platform are user-friendly,easy expansion,and can automatically retrieve accurate and valid data from large scale web pages.