Cost-Effective Hardware-Aware Machine Learning to Mitigate Memory and Workload Bottlenecks in Parallel Architectures

Pablo Sánchez Cuevas · Deposito de Investigacion Universidad de Sevilla (University of Seville) · 2026

Computer Architecture is evolving rapidly to keep pace with modern computational demands. Driven by breakthroughs in Big Data, AI, and Edge Computing, we are entering a defining era for the field. However, this progress is hindered by two critical, interdependent bottlenecks: memory and workload efficiencies. In this context, this work addresses and implements cost-effective Machine Learning methods and hardware-aware models and strategies that entail relevant advancements in both the mitigation of the Memory Wall and the optimization of workload modeling and allocation. First, the memory access patterns of state-of-the-art algorithms are analyzed, along with the memory hierarchies of modern micro-architectures, identifying the potential performance improvements provided by cache memory prefetchers. To implement an optimized framework, the Support Vector Machines For Address Prediction (SVM4AP) is proposed, which achieves high-accuracy predictions on memory accesses via its short-term online learning with low resources. As such, SVM4AP outperforms its counterparts in cost-effectiveness. Then, a novel family of cache memory prefetchers named Greedily Accurate SVM-based Prefetcher (GASP) is proposed. To take advantage of its notable prediction capabilities, GASP integrates the SVM4AP model to predict the next memory block that is read in the target cache. The results on the ChampSim simulator show that the standard GASP achieves a superior speedup to the state-of-the-art prefetchers while requiring an inferior hardware cost. To prove the physical viability of this proposal, the Spatial GASP (SGASP) is developed through a hardware implementation using FPGAs. While achieving a fully pipelined design with minimal hardware resources, this approach was successfully validated through two novel methodologies. On the one hand, its outputs were directly compared with the ones from the ChampSim implementation, showing a high matching rate. On the other, when integrated in a MicroBlaze-based system running the CoreMark-PRO benchmark, the application of the SGASP prefetcher results in a significant speedup, higher than alternative prefetching methods. Second, the impact of workload allocation on energy, throughput and runtime is studied for different parallel, heterogeneous and distributed architectures. Hardware-aware modeling and simulation plus optimization methods (via metaheuristic- and heuristic-based methods) are identified as critical techniques in two case studies. A first approach is applied to the energy-time modeling of a distributed metaheuristic model running in a heterogeneous cluster composed of CPU+GPU nodes. Specifically, a novel analytical model is proposed, where energy and runtime are accurately estimated via workload simulation. To fit the parameters of the non-linear simulation-based model, a Genetic Algorithm is successfully applied, achieving the energy-time estimations with minimal error. Another workload-based approach is KnowYourPlace: the proposal of a novel kernel and buffer placement method for the massively-parallel AI Engines architecture. In this case, a full hierarchical hardware model is proposed, allowing a detailed estimation of time costs and hardware constraints. Thus, optimal workload placement is computed through a heuristic graph-based search algorithm, generating reduced latency solutions. Therefore, KnowYourPlace successfully outperforms state-of-the-art placement methods in AI Engines, including both manual and vendor counterparts. Finally, this work provides an overall discussion of the here-proposed methods related to their promising results, including comprehensive conclusions and the proposal of future works.

Read the paper · More papers on PaperTik