Alice's Dice Expected Value Optimal Play
Expected Value for Rerolling Dice is a medium quant interview question on Games, reported to have been seen at Akuna Capital, Belvedere Trading, Citadel, Goldman Sachs, Hudson River Trading, Jane Street and Optiver.
MyQuantPartner is not affiliated with, endorsed by, or sponsored by these companies, and all trademarks belong to their respective owners.
This classic quant interview question is about optimal stopping in a simple stochastic game. You observe a random outcome, decide whether to accept it or pay a cost to try again, and aim to maximize your expected payoff. It ties together expected value, dynamic decision making, and the trade-off between current gains and future opportunities, which is central in many quant prep exercises and real trading problems.
It trains your intuition for optimal stopping with a cost, risk-reward trade-offs, and formulating a recursive value function. You must translate an intuitive game description into precise expectations, then reason about when to continue versus when to lock in a payoff. It also builds comfort with infinite-horizon setups that still have finite, well-defined values.
This matters for quant interviews because it mirrors real decisions in market making, execution, and strategy design. Interviewers use it to see whether you can structure a clean model from a verbal description, handle conditional expectations under uncertainty, and argue clearly for an optimal policy. It is especially relevant for candidates preparing for trading, strat, and research roles, where quant interviews emphasize both mathematical reasoning and clear communication.
What it tests
This class of problems is governed by the principle of optimal stopping with reset and cost: at each stage, the decision-maker compares the immediate reward to the expected value of continuing, factoring in any costs. The key structure is that the process is memoryless—each new trial is statistically identical to the last, so the optimal strategy is stationary and can be described by a threshold policy. The threshold is set where the expected gain from continuing (after subtracting the cost) no longer exceeds the immediate reward available. This balance is found by equating the expected value of stopping with the expected value of rerolling, leading to a self-consistent equation. The principle holds because, with independent trials and a fixed cost, the only relevant information is the current outcome, not the history.
Practise this question with written feedback, or hear it in a spoken mock interview.
Get started free