OpenCL GPU Optimization Through Memory Hierarchy and Coalescing
Summary
The article explains how OpenCL kernels can be optimized by accounting for GPU hardware. Using large matrix multiplication as its example, it introduces the OpenCL memory model, including global, constant, local, and private storage, and explains why local memory and registers can reduce expensive accesses to global memory. It also discusses register spilling and the tradeoff between memory placement and latency.
The main techniques include coalescing adjacent threads' global memory requests to use memory bus bandwidth efficiently, avoiding bank conflicts in local memory, and using private or local data for intermediate calculations. The article describes vectorization and alternative data layouts as further optimization steps. It reports a GPU speedup of about 200 to 1 over a sequential CPU implementation, while acknowledging that the CPU version was not highly optimized and that platform controls were limited. The results are tied to the hardware and implementation described, so they do not establish a universal speedup for trading workloads.
Key ideas
- OpenCL's abstract memory spaces map onto hardware with different access costs and performance properties.
- Coalescing neighboring threads' global memory reads can reduce wasted transfers and improve bandwidth use.
- Conflicting local memory bank accesses can serialize work, while same-address reads may be broadcast.
- Keeping intermediate values in registers or local memory can avoid slower global memory traffic, subject to register pressure.
- The reported GPU advantage is specific to a matrix multiplication example and a CPU baseline that was not highly optimized.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.