Discourse-Aware Referring Expression Generation in Visually Grounded Dialogue with Large Language Models: An Evaluation with Mistral
Iskander Grissel, Mateo Hollambica, Tariq McAllister, Naomi Wells · 2024
The growing demand for systems that can generate coherent and contextually accurate referring expressions in visually grounded dialogue has led to the development of more sophisticated multimodal models capable of integrating both visual and linguistic inputs. Introducing a discourse-aware comprehension mechanism into Mistral has provided a novel approach to maintaining conversational context, enabling the model to generate referring expressions that adapt dynamically to evolving dialogues while accurately grounding them in the corresponding visual elements. Fine-tuning Mistral on a range of public datasets has enhanced its ability to manage complex visual scenes, particularly in cases where multiple objects share similar attributes. The evaluation, which employed automated metrics such as BLEU, ROUGE, CIDEr, and IoU, demonstrated that Mistral outperformed baseline models in terms of both coherence and accuracy. Moreover, the model showed a high degree of computational efficiency, making it suitable for real-time applications. Although some limitations in handling ambiguous inputs were observed, Mistral’s architecture exhibited strong adaptability and precision in multimodal tasks, reinforcing its effectiveness in referring expression generation.