Using topic modeling to restructure the archive system of the German Waterways and Shipping Administration

A. Hoffmann, Minghao Shi, U. Rüppel · 2021

The German Federal Waterways Engineering and Research Institute (BAW) is responsible for a large number of technical documents in its archive system. These include the design process in accordance with VV-WSV 2107, which covers the entire planning cycle from basic evaluation to implementation planning. In the process of planning, construction and operation of objects of the hydraulic engineering infrastructure, a large and varied number of documents is being accumulated at the responsible authorities. Hierarchical filing systems provided with metadata are often not sufficient to search the documents in a targeted manner. The object of research is therefore machine learning methods that generate new classification systems on the basis of the given document stock and can integrate the existing documents into them. The filing is object-related and the clerk specifies various descriptive attributes. Of interest are now procedures that automatically generate topic models on the basis of the specified texts in the metadata documents in order to assign the documents to them. For this study, the words in the metadata attributes were combined into so-called bag of words and latent Dirichlet allocation (LDA) was applied to automatically find word groups that belong together. With the topic models generated in this way, documents can be searched according to topic composition or, in the case of a keyword search, documents can be displayed which do not contain the keyword but which match the topic. Due to the high number of topics that overlapped within the planning data and the few words per document, the algorithm found it difficult to generate unambiguous topics that could be easily interpreted by humans. In order to generate such topics, so-called Seeded LDA was used. Here the generation of topics can be influenced by setting seed words per topic. With Seeded LDA it is possible to fix certain topics while the algorithm decides others freely and finds new topics.

Read the paper · More papers on PaperTik