A Multimodal LLM-based Assistant for User-Centric Interactive Machine Learning

Wataru Kawabe, Yusuke Sugano · 2024

Teaser Figure User !MLLM I want a model that classifies various objects… Let's organize the points of discussion.What are you imagining? "Chairs, scissors, or TVs…?Do you have a suggestion?Then, how about classifying them based on how they are used?"Handheld," "Sit on," "No contact"... # Indeed!Let me try that one!A B User MLLM-based Assistant I want a model that classifies various objects… Let's organize the points of discussion.What are you imagining?Chairs, scissors, or TVs…?Do you have a suggestion?Then, how about classifying them based on how they are used?"Handheld," "Sit on," "No contact"... Indeed!Let me try that one! Figure 1: (A) Our system aims to assist users without a technical background in appropriately formulating tasks and creating training data in machine learning prototyping.(B) The system uses a multimodal large language model (MLLM)-based assistant to elicit user needs and guide them interactively toward appropriate training data.

Read the paper · More papers on PaperTik