Skip to content
All library documents

Configuring Parallelism in Numeric Libraries for Financial Data

Article Quant Q&A · Author: qarabala

Summary

The document considers whether a numeric library should parallelize operations by default, using large financial data workloads as its setting. Its answer recommends making thread use configurable at runtime and through environment variables, while choosing a sensible default. This lets colleagues use the library without having to tune parallelism for every operation, yet allows control when their workload or machine requires it.

The example from an R data-processing library illustrates that defaults depend on context: using many threads may suit a standalone analysis on one machine, while nested parallel work can cause contention. The example detects when a process has been forked and disables its own parallelism, since a higher level may already be distributing work. This is practical design guidance rather than a benchmark or universal rule; the document offers no fixed processor count or workload threshold, and performance will depend on the surrounding execution setup.

Key ideas

  • Threading can be exposed through runtime settings and environment variables.
  • A library should provide a default that users can adjust for their workload.
  • Nested parallelism can oversubscribe processors and create contention.
  • A library can reduce contention by detecting when its process has been forked.

Tags

Full text
# Practice of using parallel programming in numeric libraries


# Practice of using parallel programming in numeric libraries












This is a soft, and probably also an opinion based question.

Suppose I am writing a library for numeric linear algebra for the purposes of working with financial data. My goals are:

- Since I work with big data, I want to make it as fast as possible by using parallel computing.

- I want my colleagues to use this library without understanding/tweaking the parallelism underneath the functions.

The resulting questions are:

- Is this a good idea to make parallelism the default behaviour? For instance `sum(vector)` to parallelise the summation without asking the user.

- If yes, is there any rule of thumb that would cover the default behaviour of the task split between the processors? I.e. how many processors I should use?

Asking on QuantFinance since I am especially interested on how people in the industry tackle this. Thanks in advance.

## Answer by Bob Jansen (score 2, accepted)

https://quant.stackexchange.com/a/61777

I'm a big fan of data.table. The authors of data.table put a lot of thought into performance on big data sets and thanks to its popularity and age have had a lot of feedback and gained much experience in weighing the alternatives. I would definitely recommend that you familiarize yourself with their thinking.

They allow setting the number of threads using the function `setDTthreads()`:

> Set and get number of threads to be used in ‘data.table’ functions that are parallelized with OpenMP. The number of threads is initialized when ‘data.table’ is first loaded in the R session using optional envioronment (sic, PR made) variables. Thereafter, the number of threads may be changed by calling ‘setDTthreads’. If you change an environment variable using ‘Sys.setenv’ you will need to call ‘setDTthreads’ again to reread the environment variables.

This is specific to working with an R package but I think the principle applied carries over. Make it configurable at runtime and through environment variables and set a sensible default variable. In the case of data.table the default is quite greedy which make sense for people doing analysis on their own machine. It's also smart in that it detects being forked: if it detects that the process was forked, parallelism is removed because it's likely that parallelism is applied at a higher level and trying to use all the cores in all the threads would lead to very bad contention.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.