In March 2016, a computer program called AlphaGo defeated the world champion Lee Sedol at the ancient board game of Go â four games to one. That result shocked even its creators.
Go had long been considered AI's last frontier in board games. Chess had fallen to Deep Blue in 1997, but Go's 19Ă19 board has roughly legal positions â more than atoms in the observable universe. Brute-force search is hopeless; expert evaluation is hard to encode; and the game's patterns defy simple rules.
AlphaGo's breakthrough came from combining two ideas: Monte Carlo tree search (MCTS), which samples promising lines of play by rolling games out to completion, and deep neural networks trained not on human games alone, but on millions of games the system played against itself. Self-play generated its own data, self-corrected its own mistakes, and gradually built an intuition that surpassed humanity.
This article explores how MCTS works, how self-play acts as an engine of improvement, and what the technique has unleashed beyond Go.
Comments
Loading comments...