An NLP approach to Image Analysis
Guillermo Martinez · 2022
In Natural Language Processing, measuring word frequency combined with word distribution can yield a precise indicator of lexical relevance, a measure of great value in the context of Information Retrieval. Such detection of keywords exploits the structural properties of text as revealed notably by Zipf’s Law which describes frequency distribution as a ‘long tailed’ phenomenon. Can such properties be found in images? If so, can they serve to distinguish high content items (particular colours coded as RGBs) from low information items? To explore this possibility, we have applied NLP algorithms to a corpus of satellite images in order to extract a number of linguistic-type features in bitmaps so as to augment the original corpus with distributional information regarding its RGBs and observe if this addition improves accuracy throughout a Machine Learning pipeline tested with several Transfer Learning models.