Investigating the Generative-AI Evaluation Methods and Correlation with Fashion Designers

Hsi Yeh Wang, Surapong Utama · 2023

AI drawing tools set off a revolutionary trend in the field of image creation. However, there is still no clear and appropriate evaluation standard to rank AI graphics in fashion. In addition, most fashion industry insiders never used AI tools before. This research aims to evaluate whether AI-generated images could satisfy fashion designers’ needs by comparing automatic and human evaluations. Therefore, AI-generated fashion datasets with 25 images using Leonardo AI were created, and a survey was conducted to check how the experts ranked the AI images. Automatic evaluation methods, such as FID and Clip scores of each picture were measured to observe the correlation with human evaluation. The result showed the correlation coefficient between expert scores to FID scores is only 0.30, while the correlation coefficient between expert scores to Clip scores is 0.05. In other words, human evaluation and automatic evaluation are not so related and both have insufficiencies. Automatic evaluation is unable to provide judgments on fashion and aesthetics. The evaluations of different experts vary greatly due to the subjective consciousness and cannot provide fair and objective standards. Thus, it is necessary to create a new evaluation method that can evaluate the generated image in both fashion and AI aspects.

Read the paper · More papers on PaperTik