Text Mining Based on the Prototype Matching Method
Antonina Kloptchenko · 2003
s are submitted earlier to the HICSS minitrack chair. Upon arrival of the abstracts, minitrack chairs made decisions regarding the relevance of the prospective paper to his/her minitrack. Second, I considered full-versions of the papers for analysis. This two-level approach of organizing papers by content repeats the paper routing and allocating processes that the conference organizers and minitrack chairs follow in the determining whether a certain paper belongs to a certain theme or track. The conference organizers and minitrack chairs judge the appropriateness and relevance of a paper by establishing the relevance of its abstract to a track or a theme. 8.2 Abstract-level Analysis For the pilot explorative study, the abstracts from the entire HICSS-34 conference proceedings database were chosen. Abstracts are designed to condense research for the public eye by offering a preliminary overview of the research in a brief form (dos Santos 1996). At HICSS, track and minitrack chairs collect the abstract versions of the papers to be submitted to their tracks two months prior to full paper submission. Moreover, early submission of abstracts gives the chairs the opportunity to suitably allocate reviewers for the full papers. Based upon the content of the abstracts, they either recommend the authors to submit the full-paper to this particular minitrack or to look for an alternative track/minitrack. Several separate experiments were conducted to test the ability of the proposed prototype matching method to retrieve papers that are the most similar in meaning from the scientific conference collection. From every abstract I omitted the abstract titles, and author listing as irrelevant and keywords as redundant information. First, I examined the system’s ability to retrieve the most similar abstracts from the entire conference collection using any chosen abstract as a prototype query to cluster the collection. The proximity tables were constructed, where the abstract-prototype papers that were semantically closest appeared on the top, and the least similar one closer to the bottom of the table. The abstracts from the top of the proximity table were inspected. Because conference tracks are meant to unite papers from the same research field, the majority of the closest matches to every prototype were assumed to be from the same track (“track” experiment). Second, the consistency of the cross-track themes proposed by the conference organizers was analyzed. Because themes are supposed to unite the papers from different tracks that are semantically similar, I expected that abstracts from the same theme but different tracks would appear as the closest matches to an abstract from a given theme (“theme” experiment). A detailed description of the applied methodology of prototype-matching for IR by content from the HICSS-34 paper collection, the word and sentence histogram creation process, the experiments, and the line of reasoning can be found in (Kloptchenko, Back et al. 2002) and, more briefly in Paper 4. The main results and conclusions are presented below. Table 8-1 contains the results from the first “Track” experiment in the form of hit ratios per track (hit ratio 1 and hit ratio 2), which reflect how many abstracts from the same track were retrieved among the 47 or 25 closest matches on the sentence level. It is believed that the sentence level’s clustering conveys a higher