NTT’s LLM “tsuzumi”: Capable of Comprehending Graphical Documents

Ryota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida · NTT technical review · 2024

Large language models (LLMs) are being applied to fields such as healthcare, customer support, and office digital transformation.Information handled in such fields includes not only text but also a variety of visual content such as figures and diagrams.To develop LLMs as the core of artificial intelligence, their capabilities must be expanded so they can comprehend visual information.NTT Human Informatics Laboratories has been researching and developing NTT's LLM called "tsuzumi."In this article, we discuss our efforts related to tsuzumi's visual machine reading comprehension technology for comprehending the content of a document from visual information.

Read the paper · More papers on PaperTik