Distributed service-oriented architecture for information extraction system "Semanta"

L. Jastrzebski, Maciej Piasecki, G. Strzelecki, Kazimierz Wilkosz · 2005

Our objective is to provide a flexible, scalable, distributed architecture that assures a high performance for information extraction (IE) systems working in Internet. The architecture is based on both the general paradigm of the service-oriented architecture, client-server approach and strong separation of concerns between storage and processing components. An experimental IE system, named Semanta, utilising the proposed architecture is also presented. In the following document, we describe five main Semanta services, which are Web user interface (WebUI), Web crawler service (WCS), parsing service (PS), IE service and manager

Read the paper · More papers on PaperTik