Next-gen Infrastructure for Scalable Generative AI: Focus on Advances in Storage, Computing and Orchestration
Robert C. Haas · 2025
The rapid rise of generative AI is driving a shift in focus from accelerators alone to the broader infrastructure stack, with storage emerging as a critical component. Beyond scaling capacity, storage is now pivotal in enabling offload opportunities that enhance performance and efficiency. This talk explores the resurgence of interest in magnetic tape-based storage and the growing adoption of computational storage to support demanding workloads such as retrieval-augmented generation and advanced optimizations of KV caching. I will also discuss the orchestration challenges of inference across distributed infrastructures, highlighting solutions such as the opensource vLLM inference server and the llm-d distributed serving stack. Finally, I will introduce inmemory computing as a disruptive innovation in classical computing and share our latest inference results using transformer models within this paradigm.