Co-Cleansing of Image and Text Dataset by Automatic Image Annotation and Proper Noun Analysis

Yasuhide Mori, Shinrichi Satoh · 2018

There are many types of relations between images and texts in native multimedia data on the Web describing, for instance, general subjects, events, persons, etc. We propose a novel automatic multimedia data cleansing method that selects pairs of an image and a piece of text (paragraph or sentence) describing only general subjects. The selected dataset describing the general subject is thus available to be used for annotation learning. Experiments conducted on Wikipedia data confirmed that our method can automatically select image and text pairs describing general subjects.

Read the paper · More papers on PaperTik