Ontology-driven Information Retrieval in FF-Poirot.
Roberto Basili, Marco Cammisa, Maria Vittoria Marabello, Marco Pennacchiotti, Dario Saracino, Fabio Massimo Zanzotto · 2005
Abstract — This paper proposes a new approach for supporting domain information retrieval and information extraction on the web, using an original query expansion technique supported by an ad-hoc ontology focused on a specific domain of interest. The system has been built and tested in the framework of the FF-Poirot project, for supporting fine-grain retrieval from the Internet aiming at detecting financial fraudent sites. In a first stage, using a short list of keywords given by the user, the application mines the web retrieving relevant documents. These documents are then clustered into coherent groups focusing on specific subjects. The ontology model is devoted to represent the most important concepts of the domain of interest and to link them to the user need as expressed by the keywords. Once clusters of documents are made available after the first stage, the ontology can be used to extract from these clusters the most interesting documents (the most probable fraudolent sites in the framework of the FF-Poirot application). Browsing the ontology and selecting specific concepts, the user starts a query expansion engine that refines the search, creating a new query based on terminological evidences tied in the ontology to the selected concepts. The paper describes the overall software architecture of the application as used in the project, focusing specifically on the query exapansion engine and the supporting ontological model adopted. Experimental evidences, as emerged in FF-Poirot, will be used to prove the feasibility and the advantages of the adopted technique. I.