Unsupervised Keyword Extraction for Japanese Legal Documents

Tho Thi Ngoc Le, Le-Minh Nguyen, Akira Shimazu · Frontiers in artificial intelligence and applications · 2013

This study proposes a novel unsupervised approach for extracting keywords from Japanese legal documents by applying knowledge of Japanese syntax. Japanese keywords usually occur in chunks; the task of extracting Japanese keywords is treated as a matter of finding chunks that yield documents' important content. To find these chunks, all chunks in a given document are assigned weights to indicate their importance. Highly weighted chunks are recognized as candidate keywords, which are post-processed to obtain keywords. Although the proposed method employs simple techniques, the experimental results on Japanese legal documents show that the proposed chunk-based approach achieves better performance (10.5% higher on F1-score) than the graph-based ranking approach, the most popular unsupervised method.

Read the paper · More papers on PaperTik