Sequential Decisions with Noisy Measurements and Changing Payoffs
Summary
The document poses a decision problem in which a player observes a noisy measurement of an unknown payoff and must decide whether to accept it. It asks how the decision changes when a second player receives an independent noisy measurement and can claim the payoff first. The author considers whether observing the other player’s choice provides information about the underlying value, and sketches a conditional expectation that would account for that selection.
Two extensions add a payoff that moves randomly over time, with the players able to observe its direction of movement but unable to measure its level again. The final case combines this evolving payoff with competition between the players. The text offers no solution, derivation, or numerical evidence; it is an open problem prompt rather than a worked strategy. Its useful themes are Bayesian inference from another agent’s action, optimal stopping, and the effect of competition and information on accept-or-wait decisions.
Key ideas
- The initial decision is whether to accept a payoff based on a noisy observation of its unknown value.
- A rival’s decision to claim or pass can convey information about the payoff because their measurement is also noisy.
- The conditional expectation must account for both players’ observations and the selection caused by who acts first.
- Random changes in the payoff create an optimal stopping problem even when its level cannot be remeasured.
- The document poses these extensions but does not solve them.
Tags
Full text
# Game Theory Brainteaser # Game Theory Brainteaser Seeking help / thought process guidance on the following interview problem, which seems centred on game theory Setup: there’s a number X which we can measure once with error following N(0, 1). We can choose whether to receive $X (negative means lose money) a) When do we choose to receive X? b) Now another agent can also measure X with iid error, and choose to receive X before you (if they do so we can no longer receive X). What’s the strategy now? c) Back to single-player. Now after measuring, X moves up/down by one every second with even probability. We cannot observe again but know each second which direction X moves. What’s the strategy? d) Same as c) except the other agent comes back in and has the same information/setup (iid measure, can see X’s moves, can collect before us each second). What’s the best strategy now? For a), I was thinking we should take the observed value as the best estimate and accept if it is positive. For b), I guess if the other player doesn’t accept when we see a positive value, that makes it more likely the true value is lower? For (b), suppose we observe y_1 = X + e_1, and the other agent observes y_2 = X + e_2. Then I guess we want to compute E(X | y2 < 0) = y_1 - E(e_1 | y_1 - e_1 + e_2 < 0)? I tried bashing it out, but this seems to require some convolution integral, or likely that I'm overcomplicating it...
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.