Entropy-Regularized Mean-Variance Portfolios in an Asymmetric-Information Game
Summary
This paper studies two investors who choose mean-variance portfolios while competing over relative wealth. One investor knows the true stock dynamics; the other must infer them from observed market evolution. Their decisions are linked because each investor evaluates terminal wealth against the pair’s average, and the informed investor acts as leader while the partially informed investor responds as follower.
To limit information leakage, the leader randomizes her actions using an entropy-regularized objective. The follower sees realized trades but not the strategy that generated them, so her objective depends on the observed path. In the idealized continuous-observation setting, the paper derives an equilibrium with a linear follower response and Gaussian leader actions. With discrete observations, it establishes an approximate, ε-equilibrium. The account is theoretical: it reports equilibrium properties rather than empirical portfolio performance, and the continuous-sampling result may not directly represent practical trading conditions.
Key ideas
- Relative wealth concerns connect the two investors’ mean-variance decisions.
- The partially informed investor acts as follower and responds to the leader’s observed trades.
- Entropy regularization makes the informed leader’s strategy randomized to reduce information leakage.
- Continuous observation yields a linear follower response and Gaussian leader actions.
- Discrete sampling supports an approximate Stackelberg equilibrium.
Tags
Full text
# Mean-Variance Stackelberg Games with Asymmetric Information # Mean-Variance Stackelberg Games with Asymmetric Information This paper considers two investors who perform mean-variance portfolio selection with asymmetric information: one knows the true stock dynamics, while the other has to infer the true dynamics from observed stock evolution. Their portfolio selection is interconnected through relative performance concerns, i.e., each investor is concerned about not only her terminal wealth, but how it compares to the average terminal wealth of both investors. We model this as Stackelberg competition: the partially-informed investor (the "follower") observes the trading behavior of the fully-informed investor (the "leader") and decides her trading strategy accordingly; the leader, anticipating the follower's response, in turn selects a trading strategy that best suits her objective. To prevent information leakage, the leader adopts a randomized strategy selected under an entropy-regularized mean-variance objective, where the entropy regularizer quantifies the randomness of a chosen strategy. The follower, on the other hand, observes only the actual trading actions of the leader (sampled from the randomized strategy), but not the randomized strategy itself. Her mean-variance objective is thus a random field, in the form of an expectation conditioned on a realized path of the leader's trading actions. In the idealized case of continuous sampling of the leader's trading actions, we derive a Stackelberg equilibrium where the follower's trading strategy depends linearly on the actual trading actions of the leader and the leader samples her trading actions from Gaussian distributions. In the realistic case of discrete sampling of the leader's trading actions, the above becomes an $ε$-Stackelberg equilibrium.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.