Data landscape for multilingual AI
Peng Edward Wang, Pete Smith · 2025
To create multilingual artificial intelligence that is on par with the human mind, linguistic data, including text corpora and lexical resources, serves as both the access point and raw materials for computer programs to play with. By converting the same input to context-specific data representations, a system extracts relevant insights corresponding to the problem space at hand. Data is the spirit of the communication framework of multilingual AI, including its mathematical aspect, as well as the other side of the coin, the information theory. Basically, an AI-related process can be organized in two ways: around its workflow actions (what is happening) or around its data (what is being manipulated). While the workflow approach prioritizes human understanding and learning, it is intertwined with data. A well-designed workflow contains operations that are necessary not only to cultivate relevant data, but also facilitate machine learning on the fly. Linguistic resources produced at each stage of a data flow can be consumed by both humans and machines. In a data-driven process, raw materials are refined for the purposes of pattern recognition through relevant stages of Exploratory Data Analysis.