Evaluation and Benchmarking the Agent Marketing Dialogue Scenarios with Large Language Models
Kexin Zhao, Jinan Xu, Jing Shi, Yufeng Chen · 2025
In natural language processing, abilities like text comprehension, reasoning, and generation, which are usually possessed by humans, can often measure the intelligence of an artificial intelligence (AI) model. To a certain extent, measuring the model's ability to process text information can reflect the ability to use natural language. Moreover, there is a lack of exploration of the natural language level of Large Language Models (LLMs) in the actual engineering scenarios of agent marketing conversations. Therefore, we construct the dataset LURG-TEXT and its benchmark, which covers the task scenarios commonly used by natural language processing in marketing scenarios, and contains the basic indicators of the execution benchmark. Through LURG-TEXT, we conducted empirical research on some existing open-source pedestal LLMs and closedsource LLMs, focusing on the text generation and comprehension capabilities of these models in agent marketing dialogue systems. The results show that the current state-of-the-art LLMs perform well in text generation and comprehension, while the reasoning ability needs to be improved.