TERSE: A Token-Efficient Serialization Format for LLM Input
Rudson Kiyoshi Souza Carvalho · Zenodo (CERN European Organization for Nuclear Research) · 2026
TERSE (Token-Efficient Recursive Serialization Encoding) is a text-based data serialization format designed to represent the complete JSON data model with substantially fewer tokens, making it significantly more cost-efficient for use as input to Large Language Models (LLMs). Unlike existing token-efficient formats such as TOON, TERSE fully supports the entire JSON value space — including arbitrarily nested objects, heterogeneous arrays, all primitive types, and empty collections — while achieving 30–55% token reduction compared to JSON across representative document types. This document presents the TERSE format specification v0.5, including formal ABNF grammar, serialization and parsing rules, conformance requirements, security considerations, and token efficiency benchmarks. Reference implementations are available for TypeScript (terse-js) and Python (terse-py) at https://github.com/RudsonCarvalho/terse-format