Towards automatic column-based data object clustering for multilingual databases

Wael M. S. Yafooz, Siti Z. Z. Abidin, Nasiroh Omar · 2011

The amount of data in all computer applications is growing tremendously. As a result, the organization of the huge data is crucial. Recently, many researchers consider clustering as one of the important approaches in handling data for wide range of research domains. The examples include Topic Detection and Tracking (TDT), Multilingual Document Clustering, Multilingual News Clustering, Text Clustering and Web Record. Normally, data clustering is time consuming and challenging since they involve heavy programming or scripting. In online news, data clustering analysis is very much needed as the nature of the news across the globe is dynamically changing in every second. The news can come from any web sources in the form of multilingual news. This paper proposes system architecture for an automatic data object clustering in multilingual database for online news, web record and text mining. The architecture provides an overview of a virtual scheme that handles data objects within the database tables as part of the database management system. The proposed technique architecture will provide the platform for quick extraction, data arrangement, data grouping based on pattern similarities. Thus, it will improve query processing performance in multilingual databases without the need to code or script for interface programming. This is the first attempt to apply the data clustering technique prior to data extraction in any database application in the form of semi-structured and structured data (web record).

Read the paper · More papers on PaperTik