Vision-Language Navigation for Quadcopters with Conditional Transformer and Prompt-based Text Rephraser
Zhe Chen, Jiyi Li, Fumiyo Fukumoto, Peng Liu, Yoshimi Suzuki · 2023
Controlling drones with natural language instructions is an important topic in Vision-and-Language Navigation (VLN). However, previous models can not effectively guide drones with the integration of multimodal features, as few of them exploit the correlations between instructions and the environmental contexts and consider the model’s capacity to understand natural languages. Therefore, we propose a novel language-enhanced cross-modal model that has a conditional Transformer to effectively integrate the multimodal features, i.e., the textual instructions and visual contexts. To enhance the ability of language representation, we also employ SentenceBERT. In addition, to address the issue that users could provide various textual instructions even for the same navigation task, we propose a prompt-based approach by introducing an LLM-based intermediary component (LLMIR) for rephrasing users’ instructions. We evaluate our approaches with a quadcopter simulator. Our model improves the absolute task completion rate by 1.39%. To evaluate LLMIR, we create a new test set by extracting the essential and minimal instructions from the original test set. By using the LLM, the task completion rate improves by 1.51%. And it narrows the performance gap between new and original test set by 34.83%.