Skip to content
All library documents

Specialized Hardware for Low-Latency Trading and Financial Computation

Article Quant Q&A · Author: Dmitri Nesteruk

Summary

The document surveys specialized computing used in trading, including FPGA-based ticker plants, network cards designed to reduce software and kernel overhead, and hardware for feed handling and real-time risk. It also describes GPUs for computational work with limited input and output, along with configurable processors and memory-centered architectures proposed for data-heavy financial engineering tasks. Examples distinguish systems aimed at low-latency market data processing from accelerators for pricing, risk, and numerical workloads.

The discussion suggests that programmable or reconfigurable devices are often more practical than custom ASICs when deployment volumes are small or designs may need to change. FPGAs, GPUs, digital signal processors, and coarser reconfigurable architectures offer different performance, energy, and flexibility trade-offs. The evidence is a collection of industry examples and architectural approaches rather than a comparative benchmark; actual suitability depends on workload, latency needs, development cost, and data movement.

Key ideas

  • FPGAs can handle market data feeds and risk calculations where latency matters.
  • Specialized network interfaces can reduce latency by bypassing parts of the conventional networking stack.
  • GPUs are useful for selected financial computations, especially when a workload has limited input and output.
  • Programmable hardware is often easier to adapt than an ASIC when designs or deployment needs change.
  • Hardware choice depends on workload, data movement, energy use, performance needs, and development effort.

Tags

Full text
# What kind of specialized hardware is used in trading?


# What kind of specialized hardware is used in trading?












What kind of computer hardware, in additional to the 'conventional' fare, is actually used in trading? And what languages is it typically programmed in? I'm interested in ASICs, FPGAs, that sort of thing.

## Answer by Louis Marascio (score 6, accepted)

https://quant.stackexchange.com/a/1858

Some examples:

- Exegy's ticker plant uses FPGAs and InfiniBand.

- Redline Trading's ticker plant is packaged as a PCI card and uses the IBM Cell Processor.

- SolarFlare makes a line of 1G/10G nics that are heavily used because they also ship an alternate POSIX-compatible socket API that bypasses the kernel and uses DMA for reduced latency.

There are surely usages of FPGAs and ASICs for option pricing, real-time risk, etc. These aren't nearly as prevalent as the above examples.

## Answer by Andrey Taptunov (score 8)

https://quant.stackexchange.com/a/1761

The very good description of specialized hardware in finance can be found at Cisco.com - Algo Speed High Frequency Trading Solution section.

Their High-Performance Trading Architecture (pdf) poster is just great to find out used hardware for different purposes and there are also some presentations, white papers and videos about Cisco's solutions for financial markets on this website.

## Answer by NPE (score 6)

https://quant.stackexchange.com/a/1759

The only real use of this type of hardware in trading that I've seen is the recent spate of FPGA-based risk engines and feed handlers. See this article for some pointers; googling for some obvious keywords will provide more.

Given the very small deployment volumes, it seems unlikely that anyone would be looking at ASICs for this.

## Answer by John Channing (score 4)

https://quant.stackexchange.com/a/1757

There is a very small minority of people using nvidia GPGPUs which can be programmed with the CUDA libraries. This sort of specialist hardware can be very effective at solving certain problems - mostly where you have very little I/O.

More generally, if you are interested in how people are using GPGPUs, then I recommend taking a look at this question on StackOverflow.

## Answer by Giovanni (score 1)

https://quant.stackexchange.com/a/80954

Since the time to design and manufacture application-specific integrated circuits (ASICs) can take months, or at least weeks, ASICs are not favored for financial engineering, including computational finance. Shipment duration also adds to this time frame.

Hence, any hardware acceleration for financial engineering has to be programmable, or at least reconfigurable.

For small-scale, fine-grained reconfigurable logic, field-programmable gate arrays (FPGAs) would suffice for energy efficiency, and Pareto-optimal trade-offs between performance and energy efficiency. However, for larger requirements of computational power, the computation has to be partitioned across multiple FPGA boards. Somewhat recent research contests associated with the International Symposium on Physical Design (ISPD) around 2020 demonstrate that this can be done in small research teams in research universities, even in Brazil (e.g., UFRGS) and Taiwan (e.g., National Taiwan University).

Another alternative for reconfigurable computing is coarse-grained reconfigurable architectures (CGRAs), which allow you to reconfigure the connections between functional units, such as the integer adders and multipliers, floating-point adders and multipliers, encryption/decryption circuits, and matrix/tensor multipliers and adders. When memory subsystems are incorporated into CGRA solutions, near-memory computing solutions (such as processor-in-memory) can be used, too.

An interesting intersection of CGRAs and ASICs is hardware accelerators for machine learning, numerical linear algebra, graph computing, or even ODE solvers and PDE solvers that are based on in-memory computing. Dynamic random-access memory (dynamic RAM, or DRAM) subsystems, static RAM (SRAM) subsystems, and nonvolatile RAM (NVRAM) subsystems can be modified to perform a small number of array-based computation (including 2-D array-based computation, or matrix computation). This mitigates the "memory wall," by avoiding the transfer of large blocks of data between the memory subsystems (e.g., DRAM, SRAM, or NVRAM) and the processor. For streaming applications and other data-intensive computing, this type of memory-driven computing solution is very promising. It is ASIC-like, since the types of array-based computation is limited and fixed, once additional circuits are added to the memory subsystems to make this possible. It is like CGRA, since it would effectively serve as one of multiple functional units that can be connected to other functional units in various ways (the reconfigurable aspect of CGRA).

Application-specific processors, programmable hardware accelerators, or programmable co-processors, are an option. A popular example of this is digital signal processors, or digital signal processing (DSP) processors.

Domain-specific processor architectures, or domain-specific hardware accelerators, are another option.

Popular examples are:

- graphics processors for computer graphics, image processing, video processing, and computer vision

- programmable machine learning hardware accelerators, such as Google's Tensor Processing Unit (TPU) that is co-designed with TensorFlow

An aside: An intersection between application-specific processors and FPGA platforms is application-specific instruction set processors (ASIPs). Hardware/software profiling is used to select kernels or code blocks that are performance bottlenecks. Subsequently, hardware/software partitioning implements the performance bottlenecks on the FPGA platform, and the software/program is subsequently processed for additional performance speedup. The control and data flow graph(s), CFG + DFG or the hybrid CDFG, are analyzed to find common repeating sequences of instructions. Each repeating sequence of instructions is replaced with a macro-instruction, which extends the instruction set architecture (ISA). This instruction set synthesis step implies that the compiler has to be modified, via retargetable compilers, and the processor architecture (e.g., initially starting with RISC-V ISA) also has to be modified to accommodate/support the new macro-instructions. The extensions to the ISA can also be implemented on the FPGA platform.

Each type of these hardware platforms can provide hardware acceleration of techniques for financial engineering to varying trade-offs of performance and energy efficiency.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.