Towards End-to-End I/O Analysis and Optimization of HPC Systems
Hammad Ather · Scholars' Bank (University of Oregon) · 2026
The I/O subsystem is a critical bottleneck in high-performance computing (HPC) systems, making its optimization essential for scalability and efficiency. However, the growing complexity of modern hardware and the diversity of scientific workloads have made I/O increasingly difficult to analyze and tune. Existing optimization efforts suffer from two key limitations: (1) a disconnect between I/O profiling/tracing data, the performance issues they indicate, and the actionable optimizations to address them; and (2) reliance on exhaustive parameter searches or costly machine learning models to optimize the I/O performance. This dissertation addresses these limitations through two complementary, end-to-end approaches that bridge I/O analysis and optimization. First, it introduces a scalable and interactive I/O visualization and analysis approach that bridges the gap between trace collection, diagnosis, and tuning. This approach enables users to explore large-scale I/O traces, automatically detect bottlenecks, and generate targeted recommendations by combining interactive visualizations with cross-layer profiling and source-code insights. Second, it presents a lightweight runtime workflow that predicts recurring I/O patterns, extracts access insights, and applies on-the-fly optimizations using empirically derived rules. This approach eliminates the need for prior training, profiling, or parameter exploration, achieving dynamic I/O optimization with minimal overhead. Together, these approaches enable a continuous cycle of analysis-informed optimization, advancing the state of end-to-end I/O characterization and performance improvement in HPC environments.