Picture a row of slot machines — one-armed bandits, in old casino slang. Each has its own hidden chance of paying out, and you don't know which is which. You have a fixed number of pulls. Every pull does two things at once: it earns you whatever reward comes out, and it teaches you a little about that machine's true odds.
Here is the trap. To learn which machine is best, you have to explore — spend pulls on machines that might be duds. But every pull spent exploring is a pull not spent on the machine you already believe is best, your best chance to exploit. Lean too far toward exploring and you waste pulls on losers; lean too far toward exploiting and you may crown a mediocre machine champion forever, never discovering the jackpot next to it.
That tension — explore vs. exploit — is the multi-armed bandit problem. It looks like a casino curiosity, but it is the skeleton of almost every decision you make with incomplete information.
Comments
Loading comments...