Regular Expression Rule-Based Algorithm for Multiple Documents Key Information Extraction

Ramesh Mande, Kalyan Chakravarti Yelavarti, G. JayaLakshmi · 2018 International Conference on Smart Systems and Inventive Technology (ICSSIT) · 2018

With the increasing number of documents available on the internet and in offline databases, the task of supporting users in finding appropriate information becomes very complex using predictable methods due to the effort in retrieving and ranking results. Information extraction is concerned with the position of specific items in (unstructured) textual documents. The developed data can be useful for mining methods requiring structured input data, in dissimilarity to other text mining methods that utilize a bag-of-words method. A rule-based approach using regular expressions can be used to extract required key information from a document of any format. We demonstrate the applicability and benefit of the approach with real-world application, curriculum vitae for name, phone number, email_id.

Read the paper · More papers on PaperTik