Soft Computing for Information Retrieval in the WEB.

Enrique Herrera‐Viedma, Marı́a J. Martı́n-Bautista, Sergio Guadarrama, Alejandro Sobrino, José Á. Olivas · European Society for Fuzzy Logic and Technology Conference · 2005

1. Soft Computing The term SC refers to a family of computing techniques that, when L.A. Zadeh -the father of fuzzy logicintroduced the topic, originally comprised four different partners: fuzzy logic, evolutionary computation, neural networks and probabilistic reasoning. The term SC distinguishes these techniques from hard computing that is considered less flexible and computationally demanding. The key point of the transition from hard to SC is the observation that the computational effort required by conventional computing techniques sometimes not only makes a problem intractable, but is also unnecessary as in many applications precision can be sacrificed in order to accomplish more economical, less complex and more feasible solutions. Imprecision results from our limited capability to resolve detail and encompasses the notions of partial, vague, noisy and incomplete information about the real world. In other words, it becomes not only difficult or even impossible, but also inappropriate to apply hard computing techniques when dealing with situations in which uncertainty and imprecision are involved. The guiding principle of SC is “to exploit the tolerance for imprecision, uncertainty, partial truth, and approximation to achieve tractability, robustness, low solution cost and better rapport with reality”. All the methodologies that constitute the realm of SC (the four abovementioned and some others that have been incorporated in the last few years such as rough sets or chaotic computing) are considered complementary as desirable features lacking in one approach are present in another. Hence, the SC framework is put into effect by hybrid systems combining two or more of the constituent technologies with complementary characteristics. 2. Textual Information Retrieval IR may be defined, in general, as the problem of the selection of documentary information from storage in response to search questions provided by a user. IR systems (IRSs) are a kind of information system that deal with data bases composed of information items -documents that usually consist of textual informationand process user queries trying to allow the user to access to relevant information in an appropriate time interval. An IRS is basically constituted by three main components: (1) A documentary base, which stores the documents and the representation of their information contents. It is associated with the indexer module, which automatically generates a representation for each document by extracting the document contents. Textual document representation is typically based on index terms (that can be either single terms or sequences) which are the content identifiers of the documents. (2) A query subsystem, which allows the users to formulate their queries and presents the relevant documents retrieved by the system to them. To do so, it includes a query language that collects the rules to generate legitimate queries and procedures to select the relevant documents. (3) A matching or evaluation mechanism, which evaluates the degree to which the document representations satisfy the requirements expressed in the query, the so called retrieval status value, and retrieves those documents that are judged to be relevant to it. The underlying retrieval model of most of the commercial IRSs is the Boolean one, which is a robust and well formulated model although presents some limitations. For example, it does not consider partial relevance and is not able to rank the retrieved documents by relevance. Due to this fact, some paradigms have been designed to extend this retrieval model and overcome these problems, with the vector space model being the most representative. 3. Web Retrieval Although the textual IR techniques reviewed in the previous subsection are sometimes more than thirty years old, they still constitute the base of modern Web search engines. The popularity of the Web has transformed traditional IRSs into newer and more powerful search tools for locating content on the Internet. However, there are several differences due to the special characteristics of the World Wide Web environment. As Zadeh enunciated in his foreword for F. Crestani and G. Pasi’s edited book on “Soft Computing in Information Retrieval”, the problem of searching the Web has become far more complex that it was in the past mainly due to the increase on the size of the search space by several orders of magnitude and to the multimedia nature of Web documents, being composed of more information kinds than simple plain text. The main existing differences between Web retrieval and traditional IR, highlighting the following ones: (1) The HTML-based nature of Web documents, that make them present a structure defined by the HTML tags. EUSFLAT LFA 2005

Read the paper · More papers on PaperTik