Performance measurement techniques for multi-processor computers
John W. Roberts · 1986
024fiRfi7n Roberts, John w/p0vr 4a5670 QC100 -U56 N0.85-329™vT 9 % m c e ' .. 8.1 Addresses Used by Various Applications 1 1 -iii-8.2Global Memory Activity, by Address 12 8.3 Relative Frequency of Reads and Writes 12 8.4 MemoryAllocationto Code, Data etc 13 8.5 Effect of Context Switches on Above 14 8.6 Storage Transfer Algorithms 14 8.7 Reservation of Address Ranges 14 9. CACHES AND LOCAL MEMORY 14 9.1 Overall Hit Ratio 15 9.2 Instruction and Data Hit Ratios 15 9.3 Relative Frequency of Reads and Writes9.4 Distributions of Access Latencies 9.5 Cache Invalidations Due to Accesses by Other Processors 9.6 Cache Invalidations Caused by Local Processor Access 9.7 Replacements Due to Local Processor Accesses 9.8 Amount of Cached Data Never Accessed 9.9 Frequency of Cache Invalidations Caused by Task Changes 9.10 Frequency of Cache Misses Caused by Earlier Invalidations 9.11 Frequency of Cache Misses with Invalidated, but Correct, Data 9.12 Frequency of Cache Misses Caused by Data Replacing Instructions and vv 9.13 Cache Misses Due to a Poor Replacement Algorithm 9.14 Performance in Worst-Case Situations 10.SWITCHING NETWORKS 10.1 Traffic Relative to Bandwidth for Each Path 10.2 Network Saturation vs.Time 10.3 Delays Due to Blocking, for Each Path 10.4 Intermittent Excessive Blocking -iv-10.5 Percentage of Single-and Multi-hop Traffic 27 10.6 Number of Active Paths vs.Time 28 11.BUSES 29 11.1 Utilization Compared to Bandwidth 29 11.2 Bus Saturation vs.Time 30 11.3 Bus Access Delays 31 12. QUEUES 32 12.1 Queue Lengths vs.Time 32 12.2 Queue Lengths During Benchmark Execution 33 12.3 Distribution of Queue Lengths 33 13.PROCESSES 34 13.1 Time Spent Waiting in a Job Queue 34 13.2Time Spent Waiting for Shared Resources and Sync Signals 34 14.VARIABLES 35 14.1 Relative Number of Local, Global, Variables 35 14.2 Variables Transferred Between Processes 36 14.3 Parallel Processes Sharing Each Variable 36 14.4 Placement for Efficient Caching 37 15.INSTRUCTIONS 37 15.1 Size of Each Task in Memory 37 15.2Is Shareable Code Shared?38 15.3 Execution Path Through Instructions 38 15.4 Instruction Placement for Efficient Caching 39 15.5 Instruction Execution Times 40 16.SHARED RESOURCES (may include memory) 40 16.1 Fair Sharing of Resources 40 -v-16.2Distribution of Processors Waiting 41 16.3 Complete Execution Statistics 42 16.4Bandwidth Used to Wait for Resource 43 17.OWNERSHIP TOKENS 43 17.1 Time Spent Processing Tokens 43 17.2 Effect of Token Loss 44 18. SYNCHRONIZATION 45 18.1 Interprocess/Intertask Synchronization Cost 45 18.2 Is Static Task Allocation Correct? 45 19.PRIORITY 46 19.1 Effects of Priority Scheme 46 20.FAULTS 47 20. 1 Detection and Logging of Faults 47 20.2 Response to Faults 48 21.