Statistical Analysis of Real XML Data Collections
Irena Holubová, Kamil Toman, Jaroslav Pokorný · 2006
Abstract. Recently XML has achieved the leading role among lan-guages for data representation and thus we can witness a massive boom of corresponding techniques for managing XML data. Most of the process-ing techniques however suffer from various bottlenecks worsening their time and/or space efficiency. We assume that the main reason is they con-sider XML collections too globally, involving all their possible features, although real data are often much simpler. Even though some techniques do restrict the input data, the restrictions are often unnatural. In this paper we analyze existing XML data, their structure and real complexity in particular. We have gathered more than 20GB of real XML collections and implemented a robust automatic analyzer. The analysis considers existing papers on similar topics, trying to confirm or confute their observations as well as to bring new findings. It focuses on frequent but often ignored XML items (such as mixed content or recursion) and relationship between schemes and their instances. 1