Parallelism Detection in Nested Loops

Alain Darte, Yves Robert, Frédéric Vivien · Birkhäuser Boston eBooks · 2000

Loop transformations have been shown to be useful for extracting parallelism from regular nested loops for a large class of machines, from vector machines and VLIW machines to multiprocessor architectures. Of course, each type of machine corresponds to a different optimized code; depending on the memory hierarchy of the target, the granularity of the generated code must be carefully chosen so that memory access is optimized. Fine-grain parallelism is efficient for vector machines, whereas for distributed-memory machines, coarse-grain parallelism (obtained by tiling or blocking techniques) is preferable and permits the reduction of interprocessor communication. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Read the paper · More papers on PaperTik