Data Mining and Knowledge Discovery in Real Life Applications

2009

Knowledge Discovery and Data Mining are powerful data analysis tools.Data mining (DM) is an important tool for the mission critical applications to minimize, filter, extract or transform large databases or datasets into summarized information and exploring hidden patterns in knowledge discovery (KD).Sort of the data mining algorithms can be used in different real life applications.Classification, prediction, clustering, segmentation, summarization can be done on the large amount of data by using DM methods.The rapid dissemination of these technologies calls for an urgent examination of their social impact.The terms "Knowledge Discovery" and "Data Mining" are used to describe the 'non-trivial extraction of implicit, previously unknown and potentially useful information from data.Knowledge discovery is a concept that describes the process of searching on large volumes of data for patterns that can be considered knowledge about the data.The most well-known branch of knowledge discovery is data mining.Data mining is a multidisciplinary field with many techniques.With these techniques you can create a mining model that describe the data that you will use, some of these techniques are clustering algorithms, tools for data pre-processing, classification, regression, association rules, etc.. Data Mining is a fast-growing research area, in which connections among and interactions between individuals are analyzed to understand innovation, collective decision making, problem solving, and how the structure of organizations and social networks impacts these processes.Data Mining finds several applications; for instance, in different areas like biological, medicine, industrialist, business, economic, e-commerce, security, others.This book presents four different ways of theoretical and practical advances and applications of data mining in different promising areas like Industrialist, Biological, and Social.Twenty six chapters cover different special topics with proposed novel ideas.Each chapter gives an overview of the subjects and sort of the chapters have cases with offered data mining solutions.We hope that this book will be served as a Data Mining bible to show a right way for the students, researchers and practitioners in their studies.The twenty six chapters have been classified in four corresponding parts.• Knowledge Discovery • Clustering and Classification • Challenges and Benchmarks in Data Mining • Data Mining Applications The first part contains four chapters related to Knowledge Discovery.The focus of the contributions in this part, are the fundamental concepts and tools for extraction, representation, and retrieving of Knowledge in large Data Bases.The second part contains seven chapters related to Clustering and Classification.The focus of the contributions in this part, are the techniques and tools used to made clusters and classify the data on the basis of specific characteristics. VIIIThe third part contains three chapters related to Challenges and Benchmarks in Data Mining.The focus of the contributions is the applications and instances of problems that can be challenges or benchmarks for the developed tools to realise the Data Mining.The fourth part contains twelve chapters related to Data Minig Applications.The focus of the contributions in this part, are the real applications, e applications were classified in three topics that are Social, Biological and Indistrialist, on the basis of the applications that are approached in them. Part I: Knowledge DiscoveryChapter 1 analysis SE (software engineering) process models and proposes a joint model based on two SE standards and claims that this comparison revealed that CRISP-DM does not cover many project management, organization and quality related tasks at all or at least thoroughly enough (Marbán, et.al.).Chapter 2 offers an explanatory review on mining large-scale datasets and the extraction, representation, and retrieving of knowledge on Grid systems.Different trends and a domain independent solving environment ADMIRE also presented (Aouad, et.al.).Chapter 3 focuses on the main advantages of rough set theory (RST) in data analysis.Also, mentioned that RST does not need any preliminary or additional information concerning data, such as basic probability assignment in Dempster-Shafer theory, grade of membership or the value of possibility in fuzzy set theory (Rissino and Lambert-Torres).Chapter 4 introduces a novel robust data mining method by integrating a DM method for pre-processing unclear data and finding significant factors into a multidisciplinary RD method for providing the best factor settings (Shin, et.al). Part II: Clustering and ClassificationChapter 5 presents association rule mining on the selection of meaningful association rules.As an application of semantic analysis and pattern analysis, real practice case study of traffic accidents are demonstrated (Marukatat).Chapter 6 explorers to compare hybrid cluster techniques for cognitive mapping with traditional intellectual subject-classifications schemes based on the external validation of clustering results by expert knowledge present in ISI subject categories (Janssens, et.al.).Chapter 7 describes a system to inspect electronic components and devices, specifically, LCDs Panels that are the core parts of an LCD monitor store inspection result data in the RFID TAG and the Reader/Writer for efficient production control.C4.5 algorithm and Neural Nets are applied in the manufacturing process for TFT LCDs and suggest methods by which to locate defective parts (Kim, et.al.).Chapter 8 addresses steps for processing the hyperspectral remote sensing images that atmospherically corrected before processing, and a endmember was extracted by PPI algorithm from the intersection area of multi-segmentation and geology map.Dioritic porphyrite area was extracted from hyperspectral remote sensing images by Spectral Angle Mapper (SAM), Multi Range Spectral Feature Fitting (Multi Range SFF) and Mixture Tuned Matched Filtering (MTMF) using the extracted end member respectively.Finally, classification results were outputted by combining three classification results using Classification and Regression Trees (CART) (Wen, et.al.).Chapter 9 reports content based image classification via visual learning that is a learning framework which can autonomously acquire the knowledge needed to recognize images using machine learning frameworks (Nomiya and Uehara).X set-based methods, which can be used in different contexts for understanding biological pathways and describes some case studies related with the leukemia disease data (Miyoung and Kim).Chapter 21 addresses one of the important bioinformatics issue DNA sequences and presents development of microsatellite markers by using Data Mining and authors claims that development of SSRs by data mining from sequence data is a relatively easy and costsaving strategy for any organisms with enough DNA data (Tong, et.al.). Industrialist ApplicationsChapter 22 proposes a knowledge based six-sigma model where DMAIC for six-sigma was used along with data mining techniques for identifying potential quality problems and assisting quality diagnosis in manufacturing processes (He, et.al.).Chapter 23 explores a data mining process model and appropriate data mining based decision support system to support decision processes in the direct marketing in publishing sector and claims that direct mailing process gives positive results (Rupnik and Jaklic).Chapter 24 presents a data mining based methodology to group and organize data from a dam instrumentation system aiming to assist dam safety engineers in order to select, cluster and rank 72 rods of 30 extensometers located at the F stretch of Itaipu's dam (Villwock, et.al.).Chapter 25 considers a data mining algorithm for monitoring PCB assembly quality in order to minimize visual defects (Zhang).Chapter 26 reviews data mining applications which are used in power systems.This chapter also presented results comparing different figures of merit for evaluating fault classification systems (Morais, et.al.).

Read the paper · More papers on PaperTik