Beyond Sanitizers: LLM-Guided Refinement for Differential JSON Deserialization Testing in Dart

Ziyad Mohammad Mansy Ibrahim · Zenodo (CERN European Organization for Nuclear Research) · 2026

A prior black-box, LLM-guided grammar-fuzzing pipeline raised a native C parser's acceptance rate from 58.2% to 97.1% using only coverage-free feedback, and named a memory-safe target — where "crash" has no sanitizer-based meaning — as future work. We take up that challenge, retargeting the identical refinement loop at four widely-used Dart/Flutter JSON deserialization paths (hand-written parsing, json_serializable, freezed, and built_value) and replacing the crash oracle with two that need no crash at all: cross-implementation differential divergence, and a ground-truth numeric round-trip check. In seeded n=5-per-arm comparisons, refinement reproduces the original result on a like-for-like acceptance objective (50.0% to a 92.0% mean, exact p ≈ 0.0079), but on the differential objective it finds divergence in only 2.9% of generated documents against 22.2% for a static, schema-aware generator. Lifting the sandbox's JSON-library restriction, richer feedback, incremental mutation, and a larger proposer model all fail to close this gap. Giving the proposer four observations from the thirteen-case manual characterization the static generator was designed from closes it entirely: in a pre-registered, replicated comparison, refinement then reaches a 28.8% mean (p ≈ 0.013 against score-only), matching the static generator — yet it never finds a divergence beyond what it was told. The gap is domain knowledge, not search. The campaigns also expose two practitioner-actionable defects: json_serializable and freezed silently saturate out-of-range integers to int64 bounds (5.7% of 43,336 checks), and built_value silently defaults a missing required list to empty.

Read the paper · More papers on PaperTik