Latency-Aware Write Buffer Resource Control in Multithreaded Cores
Shane Carroll, Wei-Ming Lin · International Journal of Distributed and Parallel systems · 2016
In a simultaneous multithreaded system, a core's pipeline resources are sometimes partitioned and otherwise shared amongst numerous active threads.One mutual resource is the write buffer, which acts as an intermediary between a store instruction's retirement from the pipeline and the store value being written to cache.The write buffer takes a completed store instruction from the load/store queue and eventually writes the value to the level-one data cache.Once a store is buffered with a write-allocate cache policy, the store must remain in the write buffer until its cache block is in level-one data cache.This latency may vary from as little as a single clock cycle (in the case of a level-one cache hit) to several hundred clock cycles (in the case of a cache miss).This paper shows that cache misses routinely dominate the write buffer's resources and deny cache hits from being written to memory, thereby degrading performance of simultaneous multithreaded systems.This paper proposes a technique to reduce denial of resources to cache hits by limiting the number of cache misses that may concurrently reside in the write buffer and shows that system performance can be improved by using this technique.