MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
Qinzhuo Wu, Weikai Xu, Wei Liu, Tao Tan, Liujianfeng Liujianfeng, Ang A. Li, Jian Luan, Bin Wang, Shuo Shang · 2024
Recently, mobile AI agents based on VLMs have gained increasing attention.These works typically utilize VLM pre-trained on generaldomain data as a foundation, fine-tuning it on instruction-based mobile datasets.However, the proportion of mobile UI in general pretraining data is very low.Moreover, the general pre-training task does not particularly consider the characteristics of mobile UI.Therefore, directly applying such pre-trained models for mobile UI instruction fine-tuning will not yield the desired performance.In this paper, we propose MobileVLM for Chinese UI manipulation.On top of the general pre-training model, two additional pre-training stages are implemented with four specific tasks to enhance both intraand inter-UI understanding.In addition, a large Chinese mobile UI corpus, named Mobile3M, is built from scratch to compensate for the lack of relevant data.Besides 3 million static UI pages, it also contains directed graph structures formed by real-world UI transition actions.Experimental results show MobileVLM excels on both in-house test sets and public mobile benchmarks, outperforming existing VLMs.Dataset and Code are available at https://github.com/XiaoMi/mobilevlm. * Equal contribution.