HANDLING SENTENCE COMPLEXITY IN INFORMATION EXTRACTION FOR AUTOMATED COMPLIANCE CHECKING IN CONSTRUCTION

J Zhang, Nora El-Gohary · 2013

Existing construction automated compliance checking (ACC) systems require manual effort for extracting requirements from textual regulatory documents (e.g. building codes) and encoding these requirements in a computer-processable format. To address this gap, we have proposed a new approach for ACC which automatically 1) extracts semantic information (concepts and relations) from construction regulatory documents, and 2) transforms the extracted information into Prolog logic clauses for automated reasoning. Due to the variety of natural language structures, the number of text patterns in one document could become extremely large. As a result, text processing (i.e. information extraction (IE) and information transformation (ITr)) becomes highly complex and difficult. In our approach, we are proposing two methods to handle sentence complexity: 1) topdown method: starting from the top level (i.e. full sentence) and proceeding down to identify and process complex sentence components, and 2) bottom-up method: starting from the lowest level (i.e. single terms in a sentence) and proceeding up to identify and process complex sentence components. Complex sentence components are intermediately processed segments of text that are composed of multiple concepts and relations. Further processing of complex sentence components results in recognition of concepts and relations for subsequent use in constructing logic clauses. We tested our proposed methods in processing quantitative requirements (i.e. IE and ITr) from the International Building Code. We compared the results against manually-developed gold standards; and evaluated the performance in terms of precision, recall, and F1 measure. Both methods achieved high performance, but the bottom-up method outperformed the top-down method. The bottom-up method achieved 0.962, 0.961, and 0.962, while the top-down method achieved 0.954, 0.925, and 0.939 for precision, recall, and F1 measure, respectively.

Read the paper · More papers on PaperTik