Skip to content
All library documents

Separating Policy Error from Perfect-Information Value in Event-Driven Allocation

Article Quant Q&A · Author: J.Doe

Summary

The document examines how to assess a causal portfolio controller that allocates capital among event-driven trades with different holding periods. Its root-only model optimizes expected log wealth over sampled scenarios, then follows a fixed base policy for simulated future decisions. Comparisons show higher validation multipliers than a myopic allocator, but results vary across training windows. Several proposed corrections fail to improve matched results, and analysis finds that they sometimes cut winning overnight positions. The frictionless compounding results are explicitly not live-performance evidence.

The proposed framework separates the non-anticipative optimum from a clairvoyant benchmark. A scenario-by-scenario oracle that changes the feasible actions can overstate the value of information, so the comparison should preserve the causal controller's objective and constraints. A feasible causal policy supplies a lower bound; an information-relaxation dual with a martingale penalty can supply an upper bound on the optimal causal value. Approximate dynamic programming can help improve the policy and construct penalties, but the resulting bounds depend on model quality and remain estimates rather than a direct identification of the exact optimum.

Key ideas

  • The gap between a causal policy and a clairvoyant oracle combines policy error with the value of information.
  • A fair wait-and-see benchmark should preserve the causal problem's objective and feasible action set.
  • Any feasible non-anticipative policy provides a lower bound on the optimal causal value.
  • Information-relaxation penalties based on martingale increments can tighten an upper bound.
  • Approximate value functions can support both better causal decisions and dual-bound construction.

Tags

Full text
# How can I separate MPC policy error from the value of perfect information in event-driven portfolio allocation?


# How can I separate MPC policy error from the value of perfect information in event-driven portfolio allocation?












I am building an event-driven portfolio allocator for three short-horizon trading strategies. Some positions close intraday, while others remain open overnight and temporarily lock capital that could otherwise be used for later opportunities.

This is a frictionless research experiment, not a live-performance claim. My question concerns the correct stochastic-control formulation and benchmark.

### Current causal controller

At decision time t, the controller observes:

- the models and number of signals currently present;

- their entry time;

- free cash;

- existing positions, their opening times, models and principal.

It does not observe future returns, future holding periods or future signal arrivals.

For each decision, I generate 1,500 historical scenarios over a rolling horizon of two trading sessions. Each scenario contains:

- returns and holding periods for the current signals;

- bootstrapped future signal batches;

- simulated release and reuse of capital.

The root allocation is continuous:

$$ a_{t,m}\geq 0, \qquad \sum_m a_{t,m}\leq 1, $$

with the residual retained as cash.

The controller solves

$$ a_t^{\mathrm{MPC}} = \arg\max_{a\in\Delta} \frac{1}{N} \sum_{n=1}^{N} \log W_{t+H}^{(n)} \left(a,\pi_{\mathrm{base}}\right). $$

Only the root allocation is executed, and the problem is solved again at the next real signal.

The important approximation is that future decisions inside each scenario use a frozen immediate log-optimal base policy. That policy is indexed mainly by the signal-model signature and entry-time bucket. It does not recursively solve the MPC with the simulated portfolio state.

Thus, this is a root-only MPC rollout, not a complete multistage stochastic program.

### Clairvoyant benchmark

For validation only, I also compute a dynamic-programming oracle that knows every future realized return and closing time.

For batch (i), its recursion is approximately

$$ V_i = \max \left[ V_{i+1}, \max_{j\in B_i} (1+r_{ij})V_{\tau(i,j)} \right]. $$

The oracle may wait or invest all available capital in the best individual trade in the batch. It is deliberately clairvoyant and is treated only as an upper bound.

I also calculate a weaker timing oracle that can decide whether to take a batch but must use an equal allocation within that batch.

### Data and results

The reference split contains:

- 452 training trades over 14 months;

- 162 validation trades over four months;

- three random scenario seeds.

The mean validation multipliers are:

| Policy | Multiplier |
| Immediate/myopic allocation | 6,222.81✕ |
| Root-only stochastic MPC | 9,822.36✕ |
| Clairvoyant timing oracle | 554,166.89✕ |
| Clairvoyant full-selection oracle | 3,205,082.60✕ |

The large numbers come from frictionless compounding and should only be interpreted relatively.

The MPC improves substantially over the myopic policy, but recovers only about 7.24% of the log-performance gap between the myopic policy and the full-information oracle.

Causal walk-forward results over April–July are also unstable:

| Training window | MPC versus myopic |
| 9 months | -12.52% |
| 12 months | -6.30% |
| 14 months | +26.38% |

### Corrections already tested

I tested three possible explanations using identical snapshots, folds and seeds.

- Hierarchical smoothing of sparse scenario states. I pooled neighbouring signal counts and entry times. This did not improve the matched walk-forward results.

- Training-only DP action-value guidance. A clairvoyant DP generated action-value labels using training data only. A causal ridge model predicted those values from observable features and blended them into the MPC scenario values. The reference declined from 9,822.36✕ to 9,788.61✕, with zero of three seeds improving.

- DP value only at the terminal boundary. The legacy MPC was left completely unchanged inside the horizon. A learned opportunity-cost penalty was applied once to capital still locked after the terminal boundary:

$$ \widetilde W_H = c_H+ \sum_j p_j(1+r_j) \exp\left(\lambda\rho_j\widehat y_j\right). $$

This produced 9,372.79✕, again with zero of three reference seeds improving. Only 11 of 36 walk-forward folds improved.

Post-hoc analysis showed that the learned correction repeatedly reduced allocations to winning overnight positions. The freed capital was then sometimes invested in worse intraday opportunities.

### Question

What is the principled way to determine how much of this gap is reducible policy error and how much is simply the expected value of perfect information?

More specifically, should the next benchmark/controller be formulated as:

- a genuinely multistage, non-anticipative scenario tree with a state-dependent recourse policy;

- approximate dynamic programming with a value function for free cash and the release schedule of existing positions;

- an information-relaxation or martingale-dual bound that penalizes the clairvoyant oracle;

- or another formulation appropriate for sparse, event-driven portfolio opportunities?

I am not looking for a software-library recommendation. I am trying to identify the correct non-anticipative benchmark and continuation-value formulation before changing the implementation again.

## Answer by almost_surely_ (score 2)

https://quant.stackexchange.com/a/85804

### Short version

You can't separate the two pieces because your only upper bound — the clairvoyant oracle — is the zero-penalty information relaxation, i.e. the loosest possible bound on the true non-anticipative optimum. The "expected value of perfect information" is defined against the true recourse optimum, not against your MPC, and that optimum sits strictly below the oracle. Until you replace the naive oracle with a penalized (dual) upper bound, the gap $\text{oracle}-\text{MPC}$ is $\text{EVPI}+\text{policy error}$ with no way to tell which is which.

The tool that does the separation is information-relaxation duality (Brown, Smith & Sun 2010) — the same primal–dual construction used to bound Bermudan options.

### First, repair the oracle

To measure EVPI you must relax only information, nothing else. Your "full-selection oracle" (all capital into the single best trade per batch) also relaxes the action set — it no longer respects your simplex $a\in\Delta,\ \sum_m a_{t,m}\le 1$. That inflates the bound and overstates EVPI. The correct wait-and-see benchmark solves, scenario by scenario, the same problem your causal controller solves,

$$\max_{a\in\Delta}\ \frac{1}{N}\sum_n \log W^{(n)}_{t+H}(a)\quad\text{but with the path }\omega\text{ revealed,}$$

and averages. Same objective, same constraints, only the information changed. Everything below assumes the oracle is repaired this way; otherwise "7.24% recovered" is measuring the value of information plus the value of being allowed to concentrate.

(Minor but useful: keep working in log-wealth, as you already do. Expected log terminal wealth is then an additive reward, which is exactly the setting where the duality below applies cleanly.)

### The vocabulary you're missing

This is solved-in-principle in stochastic programming. In standard notation (Birge–Louveaux):

- WS = wait-and-see = your (repaired) clairvoyant value.

- RP = recourse problem = the true non-anticipative optimum. This is what your MPC approximates, and it is generally intractable — that's the whole difficulty.

- EEV = expected result of the expected-value solution; your myopic policy is its analogue.

- $\text{EVPI}=\text{WS}-\text{RP}$ — the value of perfect information. Irreducible. No causal policy recovers any of it.

- $\text{VSS}=\text{EEV}-\text{RP}$ — the value of the stochastic solution, i.e. the value of modelling uncertainty at all. This is what your MPC is trying to capture.

So your measured gap decomposes as

$$\underbrace{\text{WS}-\text{MPC}}_{\text{what you measured}}=\underbrace{(\text{WS}-\text{RP})}_{\text{EVPI, irreducible}}+\underbrace{(\text{RP}-\text{MPC})}_{\text{reducible policy error}}.$$

You can't compute either term because you don't have RP. But you can bracket it, and that's enough.

### Bracketing RP (the actual method)

Lower bound $L\le\text{RP}$: the value of any feasible non-anticipative policy. Your MPC gives one $L$; a better causal policy gives a larger one. (This is the Longstaff–Schwartz "primal" idea.)

Upper bound $U\ge\text{RP}$: an information-relaxation dual bound. Pick a penalty $\pi(a,\omega)$ that is dual feasible — one that does not, in expectation, reward a non-anticipative policy for seeing the future, i.e. $\mathbb{E}[\pi(a,\cdot)]\le 0$ for every admissible adapted policy. Then for every such $\pi$,

$$\text{RP}\ \le\ \mathbb{E}\!\left[\max_{a}\big(\text{reward}(a,\omega)-\pi(a,\omega)\big)\right]=:U(\pi).$$

The penalty $\pi\equiv 0$ recovers WS — your current oracle, the loosest case. Penalties generated by the martingale part of any approximate value function $\tilde V$ are automatically dual feasible, and with $\tilde V=V^\star$ (the true value function) the bound is tight: $\min_\pi U(\pi)=\text{RP}$ (strong duality).

Then, all in the same units:

- reducible policy error $=\text{RP}-L\ \le\ U-L$,

- $\text{EVPI}=\text{WS}-\text{RP}\ \in\ [\,\text{WS}-U,\ \text{WS}-L\,]$.

The single number to drive down is the primal–dual gap $U-L$. With the zero penalty it equals the entire gap and tells you nothing — precisely your situation. A nonzero, martingale-generated penalty pulls $U$ below WS and the split appears. Once $U-L$ is small, most of the remaining distance to the oracle is genuinely EVPI, and no amount of policy work recovers it.

### This is the Bermudan primal–dual, relabelled

Worth seeing the parallel, because it tells you the construction is mature:

- your MPC rollout = a Longstaff–Schwartz-style primal (a suboptimal adapted policy → lower bound);

- your clairvoyant oracle = the dual with zero penalty (loose upper bound);

- the missing piece = the martingale penalty of Rogers (2002) / Haugh–Kogan (2004), made practical by Andersen–Broadie (2004).

Pricing an American option is "bound the value of an optimally-stopped, non-anticipative policy from both sides." You have the same structure, with a continuous allocation in place of a stop/continue decision.

### Why your corrections 2 and 3 backfired

Both blended clairvoyant / realized-future information into the objective (DP action-value labels; a learned opportunity-cost penalty on locked capital). Such a correction is legitimate only if it is conditionally mean-zero under the true dynamics — a martingale increment. Yours was fit to realized outcomes instead, so it doesn't tighten a bound, it biases the policy. That is exactly the symptom you reported: it systematically cut winning overnight positions and recycled the capital into worse intraday trades, because in-sample it "knew" which trades paid off. The information-relaxation penalty avoids this by construction — being a martingale increment, it has zero expectation for any non-anticipative policy, so it can only move the upper bound, never the policy's value. If you want to use your DP action values, use them to generate the dual penalty, not to reshape the scenario objective.

### Direct answer to your four options

- Multistage non-anticipative scenario tree with recourse — this is the definition of RP. The correct conceptual target, but generally intractable at your continuity/scale; don't try to solve it directly.

- ADP with a value function for free cash and the release schedule — your workhorse, for two distinct jobs: (i) a better primal policy (raises $L$), and (ii) the generator of the dual penalty (lowers $U$). On its own it does not answer the decomposition.

- Information-relaxation / martingale-dual bound penalizing the oracle — this is the direct answer. It converts your loose oracle into a tight upper bound and produces the split.

- So the formulation you want is the last two together: fit $\tilde V(\text{free cash, locked-capital release schedule, signal signature})$ by ADP; use it both as the primal policy and as the penalty generator; report $[\,L,\ U(\tilde V)\,]$. The gap is your reducible-error budget; the residual $\text{WS}-U(\tilde V)$ is a lower bound on the truly irreducible EVPI.

A final reframing: "the MPC recovers 7.24% of the gap" is not interpretable while the denominator is $\text{WS}-\text{myopic}$, because that denominator is mostly EVPI. Recompute it against $U(\tilde V)-\text{myopic}$, or the bracket $[L,U]$. It is entirely possible your MPC is already near the non-anticipative frontier and the "missing 93%" is information you can never have.

### References

- D. Brown, J. Smith, P. Sun, Information Relaxations and Duality in Stochastic Dynamic Programs, Operations Research 58(4), 2010 — the core result; start here.

- D. Brown, J. Smith, Dynamic Portfolio Optimization with Transaction Costs: Heuristics and Dual Bounds, Management Science 57(10), 2011 — closest application to your setting.

- L.C.G. Rogers, Monte Carlo Valuation of American Options, Mathematical Finance 12(3), 2002; and Pathwise Stochastic Optimal Control, SIAM J. Control Optim. 46(3), 2007 — the martingale dual and its extension from stopping to control.

- M. Haugh, L. Kogan, Pricing American Options: A Duality Approach, Operations Research 52(2), 2004.

- L. Andersen, M. Broadie, A Primal–Dual Simulation Algorithm for Pricing Multidimensional American Options, Management Science 50(9), 2004.

- V. Desai, V. Farias, C. Moallemi, Bounds for Markov Decision Processes (pathwise optimization), 2012 — practical ADP-generated penalties.

- J. Birge, F. Louveaux, Introduction to Stochastic Programming, 2nd ed., Springer, 2011 — for WS / RP / EEV / EVPI / VSS.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.