Study of a Fast Facet Mining Algorithm Using Different Types of Text
José Pérez-Carballo · Academy of Information and Management Sciences journal · 2013
ABSTRACTA fast clustering algorithm designed to mine facets from text was tested using several corpora containing different types of text including technical textbooks, a cookbook, web blogs, Wikipedia articles, and electronic mails. A is an aspect of a topic. For example the following sets would be reasonable facets in a cooking domain: ingredients (e.g. apples, cayenne pepper, chocolate), utensils (e.g. egg slicer, funnel, grater, potato masher), processes (e.g. basting, poaching, pressure cooking), etc. Several studies, have shown that interfaces that present results organized into categories or faceted hierarchies meaningful to users may help them make sense of their information problem as well as the information system itself. The algorithm studied here is based on the hypothesis that multi-word terms that appear in a similar grammatical context are likely to belong to the same facet. The results show a difference in performance of the algorithm depending on the kind of text. Text that tends to be more structured, such as a Java textbook, or Wikipedia articles, results in a larger number of the facets generated by the algorithm being judged useful by experts. Text that tends to be less structured and informal, such as blogs and email, results in less facets judged useful by experts.INTRODUCTIONThis paper describes tests of an algorithm called Fast Facet Identifier (FFID) described first in (Perez-Carballo, 2009). The tests presented here are intended to determine how well this algorithm performs with different kinds of documents.The Fast Facet Identifier (FFID) AlgorithmFFID (Fast Facet Identifier) is an algorithm designed to find facets in large collections of documents. In that paper it was reported that: a) FFID discovered sets of facets from large corpora in a short time, and b) the facets discovered by FFID could be useful when building ontologies, as well as user interfaces designed to help users browse large collections of information.Definition of a FacetA is an aspect of a topic (Anderson and Perez-Carballo, 2005). Consider the following example taken from (Perez-Carballo, 2009). The following would be reasonable sets of facets in a collection of documents containing cooking recipes. Each set has a name or label (e.g. ingredients) and members (e.g. apples, cayenne pepper, chocolate). Other facet sets in this domain would be: utensils (e.g. egg sheer, funnel, grater, potato masher), processes (e.g. basting, poaching, pressure cooking), dishes (e.g. ajiaco, potatoes, blackeyed peas, kale), herbs (e.g. basil, chicory, dill), etc.The Usefulness of FacetsTraditional library classifications have always been based on facets. Ranganathan (1892- 1972) described a facet system in the 1930s (Svenonius, 1992; Anderson and Perez-Carballo, 2005). A faceted hierarchical classification uses a set of category hierarchies (instead of only one). Each hierarchy corresponds to a different facet (dimension or property). Any topic can be described specifying each of the relevant facets. For example: look for all cooking recipes that involve facet ingredient chicken, and facet process grilled, and facet cuisine Spanish. Several researchers (English et al., 2002; Yee et al, 2003; Stoica et al., 2007, Hearst, 2006; Venkatsubramanyan & Perez-Carballo, 2007) have tested interfaces that use facets in order to support information exploration and browsing. Such tests have shown that user interfaces based on facets allow users to build much more effective queries, as well as support more effective browsing tools.FFID: A Good tool to Find FacetsFFID finds good facets and it does it fast. A good facet discovering system should be able to identify multi-word terms such as bengal potatoes and bhuna khichuri, decide that they belong in the same facet set, and determine a useful label for the set. In the case of the two previous multi-word terms a useful label could be dishes. …