R2-B2: A Metric of Synthesized Image’s Photorealism by Regression Analysis based on Recognized Objects’ Bounding Box

Shun Hattori, Kizuku Aiba, Madoka Takahara · 2022 Joint 12th International Conference on Soft Computing and Intelligent Systems and 23rd International Symposium on Advanced Intelligent Systems (SCIS&ISIS) · 2022

In recent years, a lot of researches on AI (Artificial Intelligence) for Image Synthesis and Image Generation have been being conducted actively, and state of the art GANs (Generative Adversarial Networks) for text-to-image have been able to generate precise images with high photorealism for a text-based user query (but also no-good images). However, it is pointed out that the precision of all images generated for a query has been not always enough high. Therefore, for practical usages, they are required to be re-ranked and/or filtered based on some sort of metric(s). This paper proposes a novel metric, R2-B2 (RR-BB), on photorealism, especially “size balance” (i.e., balance between in-image objects’ size), of a manually or automatically synthesized image by Regression analysis based on multiple Recognized objects’ Bounding Box, i.e., the position $(x, y)$ and size (width, height, or area) of objects recognized in the image.

Read the paper · More papers on PaperTik