An Improved Term Frequency-Inverse Document Frequency Method Solving Multi-Text Label Problem

Chi Zhao, Siyuan Wang, Xiwei Feng, Xiangli Qu, Yujie Wang, Pengcheng Hua, Lei Sun, Yue Zhang · 2022 Global Conference on Robotics, Artificial Intelligence and Information Technology (GCRAIT) · 2022

With the advent of the Internet era, people can express their likes and feelings without leaving home. A large amount of text data has also emerged, such as microblog comments, movie reviews, news reports and other data. Accurate tagging of text is a key step in quickly finding the text data that they are interested in. Since the frequency of words in the text message varies, the more frequency words may not be suitable for the text label, but fewer frequency words are better suited to the content of the text. This work considers how to avoid the interference of common words and extract tags more accurately. A new solution method called Term Frequency-Inverse Document Frequency and K-Nearest Neighbor combined to extract multi-text tags are proposed to solve it. The feasibility and effectiveness of this method are verified by an example - Douban film review.

Read the paper · More papers on PaperTik