Parsing Address Texts with Deep Learning Method
Selman Delıl, Birol Kuyumcu, Cüneyt Aksakallı, İsa Semih Akçıra · 2020
The variability of address writing forms makes it difficult to process free-text postal addresses. Since postal addresses have consecutive patterns, sequential probabilistic models such as CRF or rule-based methods are often used in the address parsing processes. On the other hand, since these methods are very sensitive to typographical errors, they fail to detect detailed patterns and require a serious expert knowledge of address writing to develop rules. It is possible to identify hidden address classes without requiring pre-processing and expert knowledge by using deep learning tools that provides successful results in solving many similar problems. In this study, a one-dimensional version of the CNN networks has been redesigned and applied for address parsing problem. According to the results obtained, 96% accuracy was obtained for the labeled data set consisting of close to 20.000 samples. The network architecture we have proposed in this study has a scalable nature that does not require any pre-post processing stages.