An AI-Powered Interactive Assistant: Integrating Multimodal Interaction for Enhanced User Experience
N B Yeswanth, S. David Samuel Azariya, Hari Hrithik R, Cibi Jegan A · 2024
This work presents a fully integrated multimodal AI-powered interactive assistant based on natural language processing (NLP), document summarisation, question answering, image generation using computer vision, and voice interaction. The assistant leverages state-of-the-art models such as LLAMA 2 for conversational interaction, BART for summarisation and DistilBERT for question answering, alongside Stable Diffusion for text-to-image generation, bringing all the pieces together seamlessly to support a broad spectrum of user needs. High accuracies, low response times, and intense user satisfaction in questionnaireing and summarising tasks are observed based on evaluation results, which further confirm the ability of this system for real-time, information-intensive applications. Further optimisation of response consistency for complex image prompts and context retention in extended conversations is challenging. This study provides important lessons for the potential of integrated multimodal AI systems to change human-computer interaction radically, and it discusses ethical implications in the construction of AI. Future work includes expanding an assistant’s capabilities for adaptive learning and introducing privacy-preserving features and memory-augmented architectures to support more complicated multi-turn interactions. Secondly, this work relates to a well-established research area in the field of intelligent digital assistants and paves the way for future developments in AI-driven, multimodal interfaces.