Asymmetric numeral systems

Jarek Duda · arXiv (Cornell University) · 2009

Abstract Digital audio systems commonly treat a baked waveform as the durable object that storage, samplers, processors, mixers, transports, and endpoints repeatedly decode, cache, transform, and rebake. This is effective and deeply optimized, but it can conflate two roles: the information needed to reproduce audio observations, and a particular sample-domain observation of that information at discrete coordinates. VOLE-Audio v1.1 strengthens the v1.0 disclosure by making the entropy-native representation layer explicit: durable audio may be represented as deterministic explanatory state plus residual or literal information not reproduced by the selected explanation, rather than as a mandatory full waveform upstream of observation. For authored sources, VOLE-Audio may preserve deterministic generator, event, graph, sampler, or transfer state directly. For sampled-origin content, an inverse procedural compiler searches bounded hypothesis families--for example repetition, periodicity, wavetables, oscillator or partial models, predictors, motifs, transforms, event/state logs, transfer operators, or learned deterministic predictors--and computes an exact residual for what the hypothesis cannot reproduce. The residual is not treated as vague error. In an exact profile, residual closure reconstructs the intrinsic canonical object before arbitrary sampler observations are applied. Literal sampled representation remains a mandatory fallback, including raw literal and entropy-coded literal forms. High-entropy, noisy, encrypted, or already-compressed sources are therefore allowed to converge gracefully toward literal or conventional lossless storage. The v1.1 representation distinguishes the deterministic hypothesis H, residual R, reversible residual symbolization Ψ(R), explicit entropy model M_e, and entropy-coded residual C_R. rANS is disclosed as one appropriate entropy-coding family, not as an audio generator, structure discoverer, or escape from Shannon limits. A representation may contain procedural state, dependencies, entropy models, residual pages, entropy-coded literal streams, parameter/event streams, checkpoints, indexes, and integrity metadata. Random access, looping, seeking, partial observation, corruption locality, and bounded endpoint work require explicit restart, index, checkpoint, or independently decodable block/page structure rather than assuming a single monolithic entropy stream is freely seekable. VOLE-Audio remains semantically independent of EntropyFS and DSFB. VOLE-Audio owns audio universes, object types, event semantics, clocks, residual closure, observation semantics, and endpoint materialization. EntropyFS may optionally persist and share canonical VOLE-Audio object graphs, entropy models, residual pages, dictionaries, checkpoints, and dependencies through content-addressed storage, but a standalone VOLE-Audio archive must remain independently materializable. DSFB or related observers may guide inverse-search budget allocation by studying residual trajectories, stability, drift, and encoded residual economics, but have zero authority over normative decode, candidate validity, exactness, fallback, or Pareto selection. At runtime, the desired path is not full residual decode followed by full-object PCM expansion. The architecture permits bounded entropy-page decoding, procedural evaluation, exact closure, sampler transforms, mixing, and final sample-code packing to be fused or tightly coupled for the requested observation. CUDA and ROCm/HIP-class systems are first implementation targets for GPU-resident descriptors, procedural state, entropy models, residual-page indexes, rANS state, voice state, and endpoint observations. Direct-materialization classes D0--D3 remain unchanged: endpoints still consume final sample-domain codes, not compressed rANS symbols, and commodity audio hardware still requires bounded physical elasticity such as DMA regions, FIFOs, registers, or equivalent state. The prior-art position is conservative. Arithmetic coding, ANS/rANS, FLAC/Rice-style residual coding, predictive and transform codecs, MPEG-4 Structured Audio, HILN, neural codecs, compressed-domain processing, tracker/register-log music representations, object audio, GPU audio, memory mapping, peer DMA, and content-addressed storage are all treated as established neighboring territory. The disclosed architecture is the composition: entropy-native procedural representation plus inverse proceduralization, exact residual closure, literal fallback, state-preserving sampler semantics, partial entropy materialization, endpoint-directed observation, optional EntropyFS persistence, and zero-authority DSFB search governance. The empirical program is Pareto-oriented. It measures standalone and shared representation bytes, entropy-model bytes, entropy-coded residual bytes, residual-page indexes, dependency closure, sample-domain exposure, host and GPU waveform residency, copy traffic, endpoint observation depth, interconnect traffic, latency distributions, deadline margin, underruns, random-access decode halo, energy, and reconstruction hashes. Baselines include raw PCM, native rANS literal coding, conventional lossless codecs, PCM-resident samplers, disk-streaming samplers, compressed-file decode paths, scalar/SIMD/GPU VOLE-Audio paths, and D1/D2 direct-output attempts where supported. The objective is not to prove that procedural representation always wins; it is to determine where storing deterministic explanation plus entropy-coded innovation changes the storage, transport, runtime, and endpoint materialization trade space. Keywords procedural audio; entropy-native audio; entropy coding; rANS; residual entropy; sampler; deterministic audio state; inverse proceduralization; residual coding; partial entropy decoding; waveform unbaking; late sample materialization; GPU audio; CUDA; ROCm; HIP; endpoint materialization; EntropyFS; DSFB; content-addressed audio; computational media.

Read the paper · More papers on PaperTik