Choosing GPUs and OpenCL Features for Quantitative Finance
Summary
The discussion considers GPU selection for parallel quantitative finance workloads, with the original question asking whether hardware should favor multiplication over addition. One answer suggests that for MATLAB or Python on a single machine, the distinction may matter less than the software ecosystem: CUDA has libraries that can ease GPU adoption, but can make code dependent on compatible NVIDIA hardware. The answer does not provide benchmarks comparing operations.
A second answer says common GPUs from NVIDIA, AMD, and Intel can handle basic vector arithmetic, and recommends checking OpenCL version support and testing code across vendors. It points to SAXPY as a simple example kernel and notes that OpenCL 1.2 may suffice for basic arithmetic. These are general recommendations rather than measured performance results; suitability depends on the workload, precision, device, and software stack. The discussion gives no systematic evidence about which cards are fastest for a particular quant workload.
Key ideas
- GPU choice depends on the workload and the libraries available for the chosen programming environment.
- CUDA can simplify GPU adoption, but may tie code to compatible NVIDIA hardware.
- OpenCL supports cross-vendor testing, though supported versions vary by device vendor.
- SAXPY is offered as a basic example for learning parallel vector arithmetic.
Tags
Full text
# Using OpenCL video cards to offload Quant Finance calculations, what features should I look for? # Using OpenCL video cards to offload Quant Finance calculations, what features should I look for? I'm benchmarking some software and am looking for cards that are better at parallel multiplication vs parallel addition. - Is there any prior work that may have this information? - What GPU features should I look for? ## Answer by zuiqo (score 1) https://quant.stackexchange.com/a/4755 That depends on your application, obviously. If you intend to run Matlab or Python on a single machine, and you're looking into which graphics card to buy, multiplication vs addition should not matter much. I that situation I would recommend an Nvidia card which features CUDA. For CUDA, there are lot of libraries available which make it easy to adapt existing code to run on the GPU. Of course you can add more GPUs for more performance using SLI whatever your Card requires. Mathworks has a nice overview that will help you getting started. For Python there is PyCUDA, but i have only very limited experience with that. For Java and C++ there are options as well, bu I've never used those. The downside of all this is that your code will be less portable as you will need to use gpuArrays (in Matlab), so if someone without a CUDA-Configuration attempts to run the code, it will fail. I have yet to find an elegant way around this (!= my boss sitting at my desk...) ## Answer by BoeroBoy (score 1) https://quant.stackexchange.com/a/57500 I've dabbled a little with this a bit. OpenCL devices should work fine - even if you use NVidia for it. I actually keep all three vendors in one machine for testing, with an NVidia, AMD, and Intel GPU. All of them are fine for basic parallel vector math for things like Quant. The difference is going to be the supported version of OpenCL. NVidia makes some of the best devices but they don't care much for OpenCL obviously as they have CUDA, so they usually only support up to OpenCL v1.2. AMD and Intel are good up to OpenCL v2.0+ and Intel will soon be releasing support for v3.0. Even your built-in Intel GPU can do that pretty well if you're just trying to learn. That said for basic arithmetic at scale, any of them will generally be fine for integer and 16/32/64 bit float arithmetic. You can also use OpenCL to accelerate min/max operations over a data set. Just check out the classic SAXPY example kernels out there. I don't think you'll need anything newer than OpenCL v1.2 which runs on all of the most common GPUs. I often run the same code on all three to test compatibility. Start with the SAXPY sample here: https://anteru.net/blog/2012/getting-started-with-opencl-part-2/ I write my own version here: https://medium.com/hashicorp-engineering/all-things-gpu-part-2-4ac1c30a20ed Hope this helps someone down the line.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.