Hive: A secure, scalable framework for distributed Ollama inference

Domen Vake, Jernej Vičič, Aleksandar Tošić · SoftwareX · 2025

Large Language Models (LLMs) require substantial computational resources, often necessitating distributed inference across multiple machines. However, organizations frequently struggle to unify scattered resources, whether due to firewalls, private networks, or heterogeneous hardware. Traditional approaches demand complex VPN configurations, direct network exposure, or manual workload balancing, making large-scale deployment impractical for many. We present Hive, an open-source framework designed to seamlessly integrate fragmented compute resources into a single, unified inference system. Hive consists of HiveCore, a central proxy that handles client requests, authentication, and task queuing, and HiveNode, a lightweight worker agent that connects securely to HiveCore and executes inference locally. By relying solely on outbound connections, Hive allows remote and isolated machines running Ollama to contribute computational capacity without requiring public network exposure. This architecture enables organizations to leverage both high-performance clusters and legacy hardware, dynamically scaling LLM inference without operational overhead.

Read the paper · More papers on PaperTik