Precise Image Editing with Multimodal Agents
Bin Fu, Chi Zhang, Fukun Yin, Cheng Pei, Zebiao Huang · 2024
The rapid advancements in large language models (LLMs) have revolutionized the field of artificial intelligence, enabling the development of sophisticated agents capable of performing complex tasks. The emergence of multimodal LLMs, such as GPT-4 Vision, has further expanded the possibilities by allowing agents to process and understand visual data directly. However, current end-to-end solutions for image content editing often fall short in terms of stability, precision, and interpretability. Motivated by these limitations, we propose a novel framework that leverages the capabilities of multimodal agents to execute precise image editing tasks in a sequential and logical manner. Our approach integrates a comprehensive suite of advanced image editing tools into the action space of the agent, enabling it to interact directly with these tools through their APIs. Additionally, we introduce a reflection mechanism that empowers the agent to critically assess its own editing output and refine its actions accordingly. Through extensive experimental validation across diverse image editing scenarios, we demonstrate the effectiveness and superiority of our approach over existing end-to-end methodologies. Our contributions include the development of a versatile and flexible multimodal agent framework for image editing, the establishment of a new benchmark for precision and reliability in AI-driven image editing, and the laying of groundwork for future research in the integration of cognitive capabilities and visual perception in AI.