Research Assistant App with Multimodal Large Language Model

J Leung, Guansen Tong · 2025

Large language models (LLMs) are useful tools for humans to more quickly evaluate and synthesize information from unstructured information sources. Large language models possess the ability to deeply and efficiently reason and comprehend unstructured text and process multiple modalities and interactively synthesize structured information from text and multimodal data in an automated fashion. Furthermore, these models have exhibited strong capabilities as a tool for automating document analysis. Therefore, large language models have seen increasing use in making workflows involving text more efficient, particularly by researchers to aid in the research process. However, there is a need for specific research applications that are guided and assisted by LLM capabilities.We describe the development and testing of a research assistant that allows users to query and summarize research papers with the assistance of multimodal LLMs. The visual-language LLM is able to extract relevant information from a variety of unstructured documents to aid the research process. The use of multimodal models is shown to be effective and helpful for our research app.

Read the paper · More papers on PaperTik