Preventing XGBoost Thread Oversubscription in Parallel Workloads
Summary
The article investigates why XGBoost training took far longer in some Python process setups than with joblib’s loky backend. Its experiments compare direct execution and several multiprocessing approaches, then identify the OpenMP thread limit as the key factor: loky workers received OMP_NUM_THREADS=1 before XGBoost was imported. Setting the variable inside a function after the library had loaded did not change the runtime’s already initialized thread configuration.
The suggested practice is to set thread limits early, before importing XGBoost, and coordinate process-level parallelism with library-level threading to avoid oversubscribing available CPUs. The report gives timing comparisons in its own environment and describes additional controls for numerical libraries. Those results are environment-specific; thread count, workload, CPU allocation, library versions, and parallel job count can all affect the best configuration. The experiments explain a performance pitfall, but do not establish one universal optimal thread setting.
Key ideas
- Set OpenMP thread limits before importing XGBoost so the runtime can read them during initialization.
- Combining multiple worker processes with many threads per process can oversubscribe CPU resources.
- The experiments attribute the observed speed gap to thread configuration rather than process start method alone.
- A thread limit of one is presented as a useful approach for process-level parallel work, not a universal optimum.
- Benchmark results depend on the machine, workload, libraries, and process count.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.