Digital Leafleting: Extracting Structured Data from Multimedia Online Flyers

Emilia Apostolova, Payam Pourashraf, Jeffrey Sack · 2015

Marketing materials such as flyers and other infographics are a vast online resource.In a number of industries, such as the commercial real estate industry, they are in fact the only authoritative source of information.Companies attempting to organize commercial real estate inventories spend a significant amount of resources on manual data entry of this information.In this work, we propose a method for extracting structured data from free-form commercial real estate flyers in PDF and HTML formats.We modeled the problem as text categorization and Named Entity Recognition (NER) tasks and applied a supervised machine learning approach (Support Vector Machines).Our dataset consists of more than 2,200 commercial real estate flyers and associated manually entered structured data, which was used to automatically create training datasets.Traditionally, text categorization and NER approaches are based on textual information only.However, information in visually rich formats such as PDF and HTML is often conveyed by a combination of textual and visual features.Large fonts, visually salient colors, and positioning often indicate the most relevant pieces of information.We applied novel features based on visual characteristics in addition to traditional text features and show that performance improved significantly for both the text categorization and NER tasks.

Read the paper · More papers on PaperTik