Information neighbourhoods for visualization and monitoring strategies: Computational aspects
C. T. J. Dodson · 2012
Aspects of modern information systems that are challenging computational and statistical analysis are dynamic complexity, high dimensionality, and inherent stochasticity. We outline the use of geometric methods to provide information neighbourhoods for visualization and monitoring of algorithms, and dynamics of stochastic behavior trajectories. Geometrization of models of real phenomena give valuable insights through features that are invariant under the choice of coordinate representation. Here we look at computational aspects relating to the study of real problems, avoiding mathematical details. Introduction Aspects of modern information systems that have come to the fore and challenged computational and statistical analysis have been dynamic complexity, high dimensionality and inherent stochasticity. Recent themes for information reuse and integration highlighted by IRI Keynote speakers have included uncertain computation [36], and anomaly detection in the context of privacy protection [3]. By 2013, the annual worldwide internet protocol (IP) traffic is predicted to be a zettabyte (270 ∼ 1021 bytes) for which 90% of consumer IP traffic and 60% of mobile IP traffic will be video. Digital cameras are merging with smart phones, and visual computing applications for computational photography and augmented reality applications are developing rapidly [23], frequently based on information geometric representation and optimization methods. Geometrization of models of real phenomena have long been known to give valuable insights because of the established value of analytic geometric features, such as a natural metric, parallelism, perpendicularity, curvature and geodesic curves that are invariant under the choice of coordinate representation. In information geometry this corresponds to the invariance of measure functions of probability distributions under changes of parameters. Interest of geometers is stimulated by novel applications because these can point to new developments in the geometrical structures available. The natural information metric provides distances between states and along state trajectories, thus facilitating optimization strategies. Information reuse and integration, utilising information theory within information geometry brings important concepts mirroring physical theory of statistical mechanics: eg. entropy (ie the ‘mean log probability density’) and its relation to maximum likelihood methods for model optimization. We outline information geometry methodology to provide natural neighbourhoods that respect the intrinsic geometry of the space of states, for visualization and monitoring of algorithms, and dynamics of stochastic behaviour trajectories. 1. Computational Information Geometry For Exponential Families Of Distributions For a random variable x ∈ Rn, a set {pθ} of probability density functions with parameters θ = {θi , i = 1, ...,n} is an exponential family if the pθ express as functions {C,F1, ...,Fn} of x ∈ Rn and a function φ of θ = {θ1, ..., θn} as: pθ(x) = e{C(x)+ ∑ i θi Fi (x)−φ(θ)}. Exponential families of probability density functions and are very important and include Gaussian and gamma distributions. They admit simple embeddings in Rn+1. The jMEF package [19], is a Java library with Matlab interface and tutorials to create, process and manage mixtures of exponential families of probability density functions: http://www.lix.polytechnique.fr/∼nielsen/MEF/