From Ops to Dev: A Dual-Layer Programmable Data Lake With Kubernetes and WebAssembly-Based Offloading

Ho Kim, Junsu Kim · IEEE Access · 2026

This paper presents a dual-layer approach that evolves the data lake from a passive repository into a platform that developers can directly program. Although modern object stores expose limited server-side operations, control typically resides with infrastructure operators, preventing data consumers, such as data scientists, from leveraging computing resources near data. This results in repeated inefficiency when large datasets are shipped across a network for processing. To address this, we design aWebAssembly-based ‘‘Connected Data Lake’’ framework that enables data consumers to deploy and execute storage applications in aWebAssembly sandbox within storage nodes, so that custom APIs can be offloaded onto storage nodes. Using WebAssembly’s portability, we abstract and integrate Linux’s io_uring-based I/O and offloadable operations to be invoked on storage nodes, delivering high-performance, low-resource I/O on general-purpose hardware. We evaluated our prototype in two complementary experiments. First, in a log filtering scenario using approximately 29.61 GiB of raw data, the proposed data operation offloading reduced the network transfer volume by 99.978% and achieved a 45.5% faster end-to-end processing time compared to the conventional compute-pulls-data model, even with WebAssembly sandbox overhead. Second, in an S3-compatible object I/O benchmark against a conventional object store, our prototype improved GET throughput by 10.1–22.2% for small-to-medium objects, PUT throughput by up to 438.6% for small objects, and STAT throughput by 15.9%, while consistently reducing latency across all PUT percentiles for objects up to 1MB and lowering CPU per operation by 24–86% for GET and 35–77% for PUT, with a 4.0–6.2×smaller memory footprint. These results show that data consumers can safely program object storage and that the architecture provides a pragmatic path to data lake implementations.

Read the paper · More papers on PaperTik