A tutorial introduction to the minimum description length principle

Peter Grünwald · arXiv (Cornell University) · 2004

Abstract This paper discloses and explores a broad camera and media architecture called VOLE-Camera. The central idea is to move the primary persisted representation of captured visual information away from an obligatory sequence of completed image frames and toward a bounded, deterministic representation of capture-side state, reusable visual structure, time evolution, procedural generators, predictive models, and exact residual information. In the strongest source-native embodiment, a camera may operate on sensor samples, sensor calibration, exposure and readout state, lens and stabilization state, inertial measurements, and other synchronized capture telemetry before those variables are discarded, baked into a conventional image pipeline, or flattened into a post-ISP frame sequence. In an inverse embodiment, an existing conventional video file may be decoded to a declared canonical observation domain and then inverse-proceduralized into the same class of representation. A native VOLE decoder may subsequently execute that representation directly into a final presentation target without requiring the original camera codec and without requiring a complete conventional raster frame as an intermediate media object. The proposal is deliberately constrained by information theory. It does not assert that arbitrary videos can be reduced losslessly to a tiny fixed-size random seed. For exact reconstruction, the disclosed representation is more accurately written as: X = M(U, Sigma, V) xor_rho R where X is the declared captured observation, U is a versioned deterministic universe, Sigma is the procedural or predictive state description, V is the requested view, M is the deterministic materializer, and R is all residual information required to close the difference exactly. When the capture is strongly determined by compact state, |R| may become small or zero. When the observation contains substantial innovation, stochastic sensor realization, unmodeled detail, or effectively incompressible content, R may dominate. A “root seed” in this paper therefore means a canonical root of the complete entropy-bearing representation, not a claim that a fixed 128- or 256-bit value can distinguish every possible long video. The phrase procedural entropy factorization is introduced here as an architectural term, not as a claim to compute a unique decomposition of Shannon entropy or Kolmogorov complexity. It denotes the operational separation of a captured observation into bounded deterministic state and reusable generative structure, plus the information that remains necessary to reproduce the declared observation. Compression may result from this factorization, but compression ratio is not treated as its defining abstraction. The paper surveys adjacent and overlapping prior art broadly: information theory and minimum-description principles; state-space estimation; procedural graphics; object-, layer-, and model-based video coding; dynamic textures; event cameras; focal-plane and in-sensor processing; computational and compressive cameras; raw sensor formats; capture telemetry and visual-inertial estimation; light fields and immersive video; neural and explicit radiance-field scene representations; modern video predictors and film-grain synthesis; and display-side damage and stream-compression systems. The paper does not assert novelty over every combination of these fields. Its purpose is to place a broad technical architecture and its variants into the public record while identifying precisely which statements are established prior art, which are demonstrated by the existing VOLE research implementation, which are proposed engineering architectures, and which remain speculation requiring experiment. Keywords VOLE-Camera; VOLE; prior art; procedural entropy factorization; sensor-native representation; computational photography; camera codec; raw video; deterministic materialization; residual coding; procedural video; capture telemetry; visual-inertial estimation; event cameras; in-sensor processing; direct materialization; video compression; media architecture

Read the paper · More papers on PaperTik