Skip to content
All library documents

Shadow Pricing for Asynchronous Kelly Capital Allocation

Article Quant Q&A · Author: J.Doe

Summary

The document considers how to size trades when signals arrive at different times, positions overlap, and committed capital cannot be reused until positions close. It proposes a bid-price approximation: subtract an opportunity cost for locking capital from the batch’s expected log-growth objective, while enforcing available-capital and stop-risk limits. A shadow price is intended to represent the value of reserving capital for future signals.

It suggests estimating that hurdle from strategy arrival rates, trade growth estimates, and holding times, and recommends keeping the policy small because the sample contains few, dependent observations. To reduce concentration from noisy estimates, the response favors a simple benchmark that allocates equal stop risk and accepts signals in arrival order, with a fixed cash reserve. It also recommends expanding-window chronological evaluation. The proposal is a heuristic rather than a derived optimal policy: it does not provide a precise formula for calibrating the shadow price, address joint return dependence in detail, or establish empirical outperformance. Its suggested benchmark and walk-forward process are useful checks, but the stated sample limitations constrain conclusions.

Key ideas

  • A capital shadow price can approximate the opportunity cost of committing funds to a position.
  • The proposed sizing objective subtracts a holding-time-scaled capital charge from expected log growth.
  • Available cash and total stop-loss risk remain explicit constraints on allocations.
  • Arrival rates and holding times can inform the hurdle applied to new signals.
  • Chronological walk-forward evaluation and an equal stop-risk benchmark help assess the heuristic.

Tags

Full text
# 85781


# How should static Kelly sizing be extended to asynchronous trading signals with capital lock-up and stochastic future arrivals?












I am designing an event-driven capital allocator for several trading strategies. I currently have three strategies, although the allocator should eventually support additional ones.

This is a mathematical modeling question, not a request for investment advice, a package recommendation, or help debugging an existing implementation.

### Available information

Each strategy generates signals at a different and uncertain frequency. For each signal $i$, I can estimate:

- a predictive distribution for its net return per dollar of notional, $R_i$, including transaction costs;

- a stop-loss loss fraction, $s_i$;

- a distribution or estimate of its holding time, $H_i$;

- the strategy that generated it;

- the historical arrival process of signals from that strategy.

I use expected log return as the geometric-growth measure:

$$ \ell_i = \mathbb{E}\left[\log(1+R_i)\right]. $$

The corresponding geometric mean return is

$$ \exp(\ell_i)-1. $$

A possible descriptive score for comparing signals with different holding periods is

$$ \rho_i = \frac{ \mathbb{E}\left[\log(1+R_i)\right] }{ \mathbb{E}[H_i] }. $$

However, I do not assume that ranking or sizing trades by $\rho_i$ is optimal. Such a ratio may be meaningful for isolated sequential opportunities, but my positions can overlap and compete for the same capital.

I also understand that a scalar expected log return is not sufficient for position sizing. The dispersion, downside tail, dependence between simultaneous trades, and estimation uncertainty of the return distributions should matter.

### Event-driven setting

Signals arrive asynchronously. Several signals can arrive at the same time, while positions opened earlier may still be active.

Capital allocated to a position remains unavailable until that position closes. Consequently, accepting a positive-expected-growth signal now has an opportunity cost: it may prevent the allocator from accepting a better signal that arrives before the first position closes.

At an event time $t$, let:

- $W_t$ be total portfolio wealth;

- $C_t$ be capital currently available for new positions;

- $\mathcal{P}_t$ be the set of open positions;

- $\mathcal{A}_t$ be the set of newly available signals;

- $x_{i,t}$ be the notional allocated to signal $i \in \mathcal{A}_t$.

The allocation vector $x_t$ simultaneously determines the three decisions I care about.

First, the amount invested now is

$$ I_t = \sum_{i \in \mathcal{A}_t} x_{i,t}. $$

Second, the amount kept available for future signals is

$$ C_t-I_t. $$

Third, the individual values $x_{i,t}$ determine how current exposure is distributed across simultaneous signals.

At a minimum, the allocations must satisfy

$$ x_{i,t} \geq 0, \qquad \sum_{i \in \mathcal{A}_t} x_{i,t} \leq C_t. $$

If $D_t$ denotes the existing stop-loss risk from open positions and $B_t$ is the permitted portfolio stop-risk budget, another constraint could be

$$ D_t + \sum_{i \in \mathcal{A}_t} s_i x_{i,t} \leq B_t. $$

Additional limits may apply to gross exposure, individual trades, and individual strategies.

When a position closes, its capital is released and its realized profit or loss changes portfolio wealth. New decisions are then made using only the information available at that time.

### Objective

The economic objective is to maximize long-run geometric growth per unit of calendar time. A possible formal objective is

$$ \sup_{\pi} \liminf_{T \to \infty} \frac{1}{T} \mathbb{E}_{\pi} \left[ \log\left(\frac{W_T}{W_0}\right) \right], $$

where $\pi$ is an allocation policy mapping the observable portfolio state and current signals to their notionals.

For a static one-period problem, let $b$ denote the fractions of wealth allocated to a vector of opportunities with joint return vector $R$. If all opportunities are known and resolve over the same period, I understand the standard log-optimal formulation to be

$$ \max_b \mathbb{E} \left[ \log\left(1+b^\top R\right) \right], $$

subject to exposure and risk constraints.

My difficulty is extending this formulation when:

- decisions occur at irregular event times;

- positions have different and uncertain durations;

- capital is temporarily locked;

- future opportunities are not yet observed;

- signal-arrival frequencies differ by strategy;

- several positions can overlap;

- the return and arrival distributions must be estimated from limited data.

### Data limitation

This is not high-frequency trading. I have roughly 18 months of observations and fewer than about 2,000 signals across all strategies.

The effective sample size is lower because signals and holding periods overlap, observations from the same strategy are dependent, and market conditions may change through time.

I have tested many allocation ideas, but including the implementation history would obscure the underlying question. I would like to reconsider the problem from first principles.

A neural network or reinforcement-learning policy would not be appropriate for this sample size. I am looking for a parsimonious and auditable solution that can be retrained chronologically as additional observations become available.

### Main question

How should the static log-optimal allocation problem be formulated and approximated in this event-driven setting so that the chosen notionals account for current expected log growth, capital lock-up, and the opportunity cost of stochastic future signals?

I suspect that the full problem could be represented as a constrained semi-Markov decision process or as a stochastic resource-allocation problem with reusable capacity. However, estimating a large state-dependent value function from fewer than 2,000 dependent signals appears unrealistic.

A useful answer would ideally describe:

- the smallest defensible state representation;

- how signal-arrival frequencies and holding-time distributions should enter the model;

- how to represent the marginal value of keeping one additional dollar available;

- a tractable low-dimensional approximation appropriate for this sample size;

- how to favor signals with stronger expected log growth without allowing estimation error to create unstable concentration;

- a chronological validation procedure and a simple benchmark that a more complex allocator should be required to outperform.

I would be particularly interested in practical experience with simplified policies such as:

- current-only fractional Kelly;

- shrinkage toward equal stop-risk allocation;

- a receding-horizon stochastic program;

- a bid-price or shadow-price approximation for available capital;

- approximate dynamic programming with a deliberately low-dimensional value function.

The central issue is not how to solve a static optimizer numerically. It is how to value currently available capital when using it changes the set of future opportunities that the portfolio will be able to accept.

References or descriptions of comparable systems used in practice would be especially helpful.

## Answer by Russlan Ramdowar (score -2)

https://quant.stackexchange.com/a/85790

TL;DR: Trying to fit a full SMDP or RL model on ~2,000 samples will just overfit to oblivion. You don't need a massive state space. You just need a shadow price (opportunity cost) for your available capital.

Here’s the most pragmatic way to hack this together using a bid-price/shadow-price approximation.

#### 1. The Core Idea: Shadow Pricing

Instead of modeling the exact combinatorial sequence of unknown future trades, assume uninvested cash has a baseline "yield" derived from the average historical performance and arrival rates of your strategies. This is your hurdle rate, $\lambda$ (expected log-growth per unit of time for free capital).

When a signal drops, it competes against $\lambda$. If it doesn't beat the shadow price adjusted for its expected holding time, you pass or size it way down.

#### 2. The Math: Modified Kelly Objective

For a batch of new signals $\mathcal{A}_t$, you want to maximize the expected log return, minus the opportunity cost of locking up that capital $x_i$ for an expected duration $\mathbb{E}[H_i]$.

Formulate the sizing at time $t$ as:

$$\max_{x} \mathbb{E} \left[ \log(1 + \sum_{i \in \mathcal{A}_t} x_i R_i) \right] - \lambda \sum_{i \in \mathcal{A}_t} x_i \mathbb{E}[H_i]$$

Subject to your core constraints:

- $\sum x_i \leq C_t$ (cannot spend more than available cash)

- $D_t + \sum s_i x_i \leq B_t$ (stop-loss risk budget)

#### 3. Estimating $\lambda$ (The Hurdle)

Keep the state representation tiny. You don't need a neural net, you just need moving averages. Calculate:

- $\mu_j$: average arrival rate of strategy $j$

- $g_j$: average expected Kelly growth per trade for strategy $j$

- $\tau_j$: average holding time for strategy $j$

$\lambda$ is effectively the time-weighted average of your historical expected returns if you just blindly took average future signals. When market regimes shift and signal density spikes, $\lambda$ scales up (cash becomes more valuable). When things are quiet, $\lambda$ drops (take whatever edge you can get).

#### 4. Beating Estimation Error

Because you only have 18 months of overlapping data, standard Kelly will aggressively concentrate your capital into whatever strategy got lucky in sample.





#### 5. Validation & Benchmarks

The "Dumb" Benchmark: Before you spin up a convex optimizer for the math above, code a basic heuristic allocator:

- Hard-code a static cash buffer (e.g., always leave 20% in $C_t$).

- Size every incoming signal to an equal stop-risk weight ($\frac{1}{N}$).

- First come, first served until cash runs out.

If your shadow-priced Kelly doesn't beat this out-of-sample, the complexity isn't worth the compute.

Chronological Validation (Walk-Forward): K-fold cross-validation is useless here because your holding periods overlap. You must use an expanding window (Walk-Forward Analysis). Train on months 1-6, simulate month 7. Train on 1-7, simulate month 8.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.