Modern alternatives to hive: A systematic review and single-node benchmark of SQL-on-Hadoop and lakehouse engines

Szoke Mark-Andor · Information Systems · 2025

This paper presents a systematic review of modern SQL-on-Hadoop and lakehouse engines as alternatives to Apache Hive and complements it with a new experimental benchmark on a single, common platform. Twelve representative engines are analyzed (selected for architectural diversity, ecosystem relevance, and evidence availability) and a compact feature taxonomy is introduced, covering vectorized execution, whole-stage code generation, LLVM/JIT, federated pushdown, cloud columnar caches, and embedded ML primitives. In addition to the qualitative synthesis, 11 engines (all except Databricks Photon) are benchmarked locally on standardized TPC-H queries at two scale factors (SF = 1 and SF = 10) using the same hardware and software environment. All engines were tested on a single node (4 CPU cores, 8 hardware threads) to provide a common, multi-threaded baseline. A multi-node (cloud) evaluation is outside the present scope and is discussed as future work in Section 6 . The results show order-of-magnitude differences between engines: native/vectorized engines (e.g., DuckDB, ClickHouse, Impala, Velox-accelerated Presto) achieve substantially lower runtimes and more efficient scale-up than JVM-only stacks (e.g., Spark SQL, Trino) on a single node. Drill failed at SF = 10 due to out-of-memory errors; Phoenix completed SF = 1 but did not complete SF = 10 on an 8 GB system (one long query aborting). Photon is discussed from the literature because it cannot be executed outside Databricks. All scripts, data, and a runnable artifact for the benchmark are open-sourced to support reproducibility and reuse. Scope: While the primary focus is distributed SQL-on-Hadoop systems, modern single-node engines (e.g., DuckDB, DataFusion) are included where they (i) implement architectural innovations relevant to distributed analytics or (ii) are used in lakehouse deployments. These are explicitly distinguished from distributed systems in the taxonomy and results.

Read the paper · More papers on PaperTik