Study on Extraction of Keywords Using TF-IDF and Text Structure of Novels
Eun-Soon You, Gun-Hee Choi, Seung-Hoon Kim · Journal of the Korea Society of Computer and Information · 2015
With the explosive growth of information about books, there is a growing number of customers who find it difficult to pick a book. Against the backdrop, the importance of a book recommendation system becomes greater, through which appropriate information about books could be offered then to encourage customers to buy a book in the end. However, existing recommendation systems based on the bibliographical information or user data reveal the reliability issue found in their recommendation results. ∙제1저자 : 유은순 ∙교신저자 : 김승훈 ∙투고일 : 2015. 2. 5, 심사일 : 2015. 2. 9, 게재확정일 : 2015. 2. 20. * 단국대학교 미디어콘텐츠연구원(Institute of Media Content) ** 단국대학교 소프트웨어학과 (Dept. of Software Science, Dankook University) *** 단국대학교 응용컴퓨터공학과(Dept. of Applied Computer Engineering, Dankook University) ※이 논문은 2014 한국지능정보시스템학회 추계학술대회에서 발표한 논문(“도서 텍스트 본문 주제어 추출 연구[14]”)을 확장한 것임 122 Journal of The Korea Society of Computer and Information February 2015 This is why it is necessary to reflect semantic information extracted from the texts of a book’s main body in a recommendation system. Accordingly, this paper suggests a method for extracting keywords from the main body of novels, as a preceding research, by using TF-IDF method as well as the text structure. To this end, the texts of 100 novels have been collected then to divide them into four structural elements of preface, dialogue, non-dialogue and closing. Then, the TF-IDF weight of each keyword has been calculated. The calculation results show that the extraction accuracy of keywords improves by 42.1% in performance when more weight is given to dialogue while including preface and closing instead of using just the main body. ▸