A Multi-source Heterogeneous Data Storage and Retrieval System for Intelligent Manufacturing
Yaning Kong, Dongmei Li, Chunshan Li, Dianhui Chu, Zekun Yao · 2021
The manufacturing industry produces massive multi-source heterogeneous data such as text, images, audio, and video in the process of design, production, sales, and service. The major problem facing manufacturing companies is how to efficiently manage and use these data resources to create value for manufacturing reproduction. Traditional data storage and retrieval systems classify heterogeneous data according to different forms or modalities and process them separately, resulting in a lack of correlation between cross-modal data (text, image, audio, and video data cannot be checked each other). It cannot support the problem of manufacturing business processes. In this article, we designed and implemented an efficient and fast cross-modal retrieval system for multi-modal industrial data such as text and pictures to realize efficient management and retrieval of multi-source heterogeneous data. Specifically, the system calculates the multi-modal content of manufacturing design, product, service, and other data as a set of unified semantic expressions and stores it in the index structure. When the user makes a query, the index system will return all the modal data related to the retrieved content. This article conducted experiments on the Flick30k data set. The experimental results show that: (1) This system can support millions of data storage and retrieval. (2) With millions of data, the system retrieval rate is in milliseconds. (3) The retrieval accuracy is higher than traditional vector retrieval methods.