Deriving the Merton Model’s Optimal Portfolio Control from the HJB
Summary
The document explains how to obtain the locally optimal portfolio control in the finite-horizon Merton expected-utility problem. In the Hamilton–Jacobi–Bellman equation, the control is chosen to maximize a function of the value function’s first and second derivatives. Applying the first-order condition to that expression yields the candidate control, which depends on those derivatives as well as the asset’s excess return and volatility.
The apparent dependence on the value function is consistent with its role: it summarizes the best achievable expected utility from a given time and wealth state. The discussion also distinguishes the full admissible control process in the original optimization from the scalar control selected over an infinitesimal time interval in the HJB equation. Dynamic programming connects these local choices into a candidate control process. The answer sketches the derivation and interpretation but does not establish the conditions under which the candidate is globally optimal or provide a worked utility-function example.
Key ideas
- The HJB control is found by maximizing the local expression in the equation.
- Differentiating that expression with respect to the control gives the first-order condition.
- The candidate control depends on the value function’s marginal utility and curvature.
- The original strategy is a process, while the HJB control is a local decision at a given state.
- Dynamic programming links local decisions to a candidate strategy over the full horizon.
Tags
Full text
# How do we solve bellman's equation in Merton's model
# How do we solve bellman's equation in Merton's model
Studying the expected utility maximization problem in Merton's model, I'm having some difficulties. Let $t$ be a starting time, $T$ the final finite Time. We define, \begin{equation} V(t,x)=\underset{\pi \in A}{\sup}\{\mathbb{E}\left[U(X_T(\pi))\right] |X_t=x\} \end{equation} Where $A$ is the set of admissible trading strategies. We can prove that $V$ satisfies (Hamilton-Jacobi-Bellman equation in Merton Model) \begin{equation} \frac{dV(t,x)}{dt} + \sup_{\pi_t} \left( \frac{dV(t,x)}{dx} x (r+ \pi_t(\mu-r)) + \frac{1}{2} \frac{d^2V(t,x)}{dx^2} x^2 \pi_t^2 \sigma^2 \right) = 0 \end{equation}
In the literature (check below for the reference), it's said that $\textit{"a candidate for the optimal control is obtained from the first-order condition}$ $\textit{for the maximum in the HJB equation is :"}$ \begin{equation} \hat{\pi}(t,x)=-\frac{\mu-r}{\sigma^2}\frac{\frac{dV}{dx}}{x\frac{d^2V}{dx^2}}(t,x) \end{equation}
At this point, I have two questions : 1/ How can we actually obtain / prove this result ? 2/ I don't understand why $V$ is actually present in $\hat{\pi}$ 's expression : By definition, isn't V a sup on all possible $\pi$ ? Shouldn't $\hat{\pi}$ depend on everything apart from $V$ ? Thanks guys !
> Huyen Pham. Optimization methods in portfolio management and option hedging
## Answer by Quantuple (score 3)
https://quant.stackexchange.com/a/33639
I'm definitely not an expert on this topic, but it seems to me that:
- $\hat{\pi}_t$ is here defined as \begin{align} \hat{\pi}_t &= \underset{\pi_t \in \Bbb{R}}{\text{argsup}} \left( \frac{dV(t,x)}{dx} x (r+ \pi_t(\mu-r)) + \frac{1}{2} \frac{d^2V(t,x)}{dx^2} x^2 \pi_t^2 \sigma^2\right) \\ &= \underset{\pi_t \in \Bbb{R}}{\text{argsup}}\, I(\pi_t) \end{align} and is hence obtained by applying the first-order optimality condition: $$ \frac{d I }{d\pi_t}(\hat{\pi_t}) = 0 $$
- The $\pi$ which appears in the definition of the optimal value function: $$ V(t,x)=\underset{\pi \in A}{\sup}\{\mathbb{E}\left[U(X_T(\pi))\right] |X_t=x\} $$ is a full-blown stochastic process $(\pi_t)_{t \in [0,T]}$. On the other hand, $\pi_t$ which appears in the HJB equation is merely a scalar: it represents the value of the optimal control that applied over the infinitesimal interval $[t,t+dt)$ (hence, "locally", $V(t,x)$ does not depend on it). This is better understood by adopting a dynamic programming view of the problem as mentioned in @Alex C's comment. The idea is then to postulate a candidate optimal control process $\hat{\pi}_t(x)$ by extending the result which has been obtained locally (cf. Bellman's principle of optimality).Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.