Skip to content
All library documents

CUDA Matrix Multiplication with Two-Dimensional Grids and Blocks

Article QuantStart

Summary

This tutorial explains how to map matrix multiplication onto CUDA threads. Each thread computes one output cell by taking the dot product of a row from the first matrix and a column from the second. The article shows how two-dimensional grid and block indices identify output rows and columns, and how matrices stored as linear arrays can be indexed by row and column. A bounds check handles dimensions that do not fill the final blocks.

The host-side example allocates and transfers arrays, launches the GPU kernel, and compares its output with a serial CPU calculation. It introduces parallel decomposition and thread organization as foundations for later computational finance work, including numerical option pricing and finite-difference methods. The example uses square matrices and is primarily educational. It points toward shared memory as a later optimization but does not explain that technique here; its displayed code should therefore be read as a basic implementation rather than a full account of efficient GPU multiplication.

Key ideas

  • Each CUDA thread can compute one output element of a matrix product independently.
  • Two-dimensional block and thread indices map naturally to matrix rows and columns.
  • A row-major linear array represents a matrix using a row offset plus a column index.
  • Bounds checks prevent extra threads in partially filled blocks from accessing invalid matrix elements.
  • Comparing GPU output with a serial CPU calculation provides a basic correctness check.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.