Computing with perceptions for the linguistic description of complex phenomena through the analysis of time series data
Alejandro Ramos-Soto, Alberto Bugarín, Senén Barro · 2014
Nowadays, new technologies allow acquiring andarchiving vast volumes of data about time-evolvingphenomena in many crucial areas such as economy,science, and industrial processes. Examples in econ-omy include the evolution of every kind of econom-ical indicators at local or global levels, like stockfunds, electricity/gas/water consumption, price of ba-sic products, etc. In science, the amount of infor-mation collected by researchers is overwhelming andever growing, including astronomical observations byradio telescopes, space probes, etc. and data collectedfrom experiments in diverse scientific fields.In order to be useful, this data must be exploitedand explained in an understandable way, reportingfacts, advice or commands to be performed that usethe available background knowledge about the phe-nomenon under study. These objectives can only beachieved by using natural language, especially if thefinal information is to be provided by a non-expert.This is clear for example in the case of financial news-papers and scientific publications, in which data is notsimply made accessible or summarized as graphicsand tables, but arguments and conclusions need to beexplained using natural language.However, there is a lack of tools and means forprocessing and interpreting all this data using com-puters. Any organization of data as provided by acomputer, either in a numerical, categorical, and/orgraphical form, is just a tool that can be employed byhuman experts to produce an explanation in naturallanguage.Understandable linguistic descriptions of phe-nomena are provided by human experts, while com-puters just provide flexibility in storing and accessingdata. In fact, it is becoming easier and easier to col-lect data, but providing a human being with expertiseon a certain domain remains difficult and expensive.This situation clearly poses a problem, since the ra-tio data/human experts is growing dramatically as aconsequence. In summary, there is a clear need forcomputational systems able to produce automaticallylinguistic descriptions of data about phenomena.More specifically, the task of generating easily un-derstandable information for people using human lan-guage has been addressed by two fields which, inde-pendently until now, have researched the processesthis task involves: the natural language generation(Reiter and Dale, 2000) and the linguistic descriptionsof data (Zadeh, 1996).The natural language generation field focuses itsefforts on automatically obtaining texts, with the pur-pose of them being as much as possible indistinguish-able from the ones created by humans. The linguisticdescriptions of data field, which originates in the softcomputing domain, provides summaries or descrip-tions from data sets using linguistic concepts whichdeal with the imprecision and ambiguity of languagethrough the use of fuzzy sets.In this context, we propose in this Ph.D. to re-search on the linguistic descriptions of data field, cov-ering a group of soft computing-based concepts andtechniques, such as linguistic variables, fuzzy opera-tors and quantification methods. For instance, usingthis kind of solutions, we can obtain quantified sen-tences such as “most of the students are good” or “Afew days with high humidity the temperature is low”.In fact, most of the approaches for building lin-guistic descriptions described in the literature makeuse of the concept of “quantified sentence”. In thissense, the linguistic description approaches make useof two different types of quantified sentences: typeI (“Q of X are A”), as in “several dogs are brown”,and type II (“Q of D are A”), as in “a few young re-searchers have published relevant papers”, where X isa finite crisp set, Q is a linguistic quantifier and A, Dare fuzzy properties defined over X (Delgado et al.,2014).Despite its formal nature and its orientation to-wards providing meaningful information from data,as of today the linguistic descriptions of data fieldhas to face several problems as a novel research do-main. First, the sole use of quantified sentences is