Designing Deep Hedging Rewards for FX Risk and Trading Costs
Summary
The document discusses reward design for reinforcement learning applied to foreign exchange hedging. A typical objective combines portfolio gains with penalties for risk, such as variation in portfolio value, and for trading costs associated with changes in the hedge ratio. The response says technical indicators and other market variables can be provided as model features; the network may then learn how they relate to hedging decisions rather than requiring each feature to appear directly in the reward.
It also notes that risk measures used in deep hedging can affect behavior. If training data contains apparent statistical arbitrage, a learned hedger may seek to exploit it instead of focusing only on risk reduction. This makes data quality important, while the choice of risk measure may help control that behavior. The document raises oil prices and macroeconomic inputs as possible context for the currency pair, but does not specify a universal reward formula or demonstrate performance for the proposed market.
Key ideas
- A common hedging reward balances portfolio gains, risk penalties, and transaction costs.
- Technical and macroeconomic variables can be supplied as model features.
- Risk measures can shape which hedging behaviors the learner favors.
- Apparent arbitrage in training data can lead a deep hedger to exploit data artifacts.
- Reward design is context dependent, and the document provides no universally validated formula.
Tags
Full text
# How to design an effective reward function for RL-based FX hedging strategy?
# How to design an effective reward function for RL-based FX hedging strategy?
I'm working on an Reinforcement learning (RL) algorithm for optimizing a foreign exchange (FX) hedging strategy, specifically for the USD/AZN (Azerbaijani manat ₼) pair in the context of Azerbaijan's economy (which is heavily dependent on oil exports). The goal is to minimize risks (e.g., volatility in profits) and maximize profits during hedging I think the common formula of reward function is $ r_t = \text{profit}_t - \lambda \cdot \text{risk}_t - c \cdot \text{transaction_cost}_t $ or another form $r_t = \Delta V_t - \lambda \cdot \text{Var}(\Delta V_t) - \kappa \cdot |\Delta \delta_t| \cdot C_t$
Where:
- $\Delta V_t$: Change in portfolio value at time $ t $, representing profit (e.g., from hedging gains or currency movements).
- $\text{Var}(\Delta V_t)$: Variance of portfolio value changes over a lookback period, measuring risk.
- $|\Delta \delta_t| \cdot C_t$: Transaction cost proportional to the change in hedge ratio ($ \Delta \delta_t $), with $ C_t $ as the cost per unit.
- $\lambda$: Risk aversion parameter (e.g., 0.1–1.0), weighting the risk penalty. $\kappa$: Cost penalty parameter, weighting transaction costs.
But how to write it by taking into account features like technical indicators (RSI, EMA, MACD), oil prices (Brent, as a key driver for AZN), inflation rates, and the exchange rate itself?
Which factors should I know to write effective reward function?
Is there are reward function which is general for hedging?
## Answer by Adam C. Jones (score 0)
https://quant.stackexchange.com/a/85183
Yes, that's a pretty typical form seen in deep hedging papers. The technical indicators can simply be passed to the neural network as features, and (in theory) it will learn the optimal hedge taking those into account.
You also often see risk measures used in the literature. But that should come with a word of warning, see Relationship Between Deep Hedging and Delta Hedging where it is shown that, if there is statistical arbitrage in your training data, a deep hedger will attempt to exploit it, so your training data needs to be good. Alternatively, note that Is the difference between deep hedging and delta hedging a statistical arbitrage? claim that this arb-taking can be avoided with a judicious choice of risk measure.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.