NTT’s LLM “tsuzumi”: Capable of Comprehending Graphical Documents
Ryota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida · NTT technical review · 2024
Large language models (LLMs) are being applied to fields such as healthcare, customer support, and office digital transformation.Information handled in such fields includes not only text but also a variety of visual content such as figures and diagrams.To develop LLMs as the core of artificial intelligence, their capabilities must be expanded so they can comprehend visual information.NTT Human Informatics Laboratories has been researching and developing NTT's LLM called "tsuzumi."In this article, we discuss our efforts related to tsuzumi's visual machine reading comprehension technology for comprehending the content of a document from visual information.