A Study on Methods for Parsing Architectural Multi-Modal Data and Extracting Modeling Parameters
Shimei Li, Weining Song, Tan Li, Nanjiang Chen, Liefa Liao, Xuejun Zhou, Fangfang Gao, Yin Runmin · Buildings · 2025
To address information isolation and incomplete parameter extraction among multi-modal data (e.g., drawings, text, and tables) in the operation and maintenance stage of buildings, this paper proposes a multi-modal data parsing, automatic parameter extraction, and standardized integration method oriented toward 3D modeling. First, by employing vector element parsing and layer semantic analysis, the method enables structured extraction of key component geometry from architectural drawings and improves modeling accuracy via spatial topological relationship analysis. Second, by combining regular expressions, a domain-specific terminology dictionary, and a BiLSTM-CRF deep learning model, the extraction accuracy of unstructured parameters from architectural texts is significantly improved. Third, a multi-scale sliding window and geometric feature analysis are used to achieve automatic detection and parameter extraction from complex nested tables. Regarding the experimental setup: the drawings consist of a large-scale collection of DXF files stratified and randomly split into train/val/test with an approximate 8:1:1 ratio; the text set includes 1550 PDF-derived specification fragments (8:1:1 split); and the tables cover typical door/window, structural, and electrical schedules (also split ~8:1:1). F1 scores use micro-F1 (instance-level aggregation), and 95% confidence intervals and their computation are described in the main text. Experimental results show that the F1 scores for wall line, wall, and column recognition reach 98.1%, 84.9%, and 92.2%, respectively, while the F1 scores for door and window recognition are 74.3% and 76.2%. For text parameter extraction, the proposed PENet model achieves a precision of 83.56% and a recall of 86.91%. For the table task, the parameter extraction recalls for doors/windows and structure are 95.0% and 96.7%, respectively. The proposed method enables efficient parameter extraction and standardization from multi-modal architectural data, demonstrates significant advantages in handling heterogeneous data and improving modeling efficiency, and provides practical technical support for the digital reconstruction and intelligent management of existing buildings.