RustZero↗
A blank slate. A board. A little search.
AlphaZero learns Breakthrough from scratch.
Loading Rust engine…
0 pliesTap a pawn, then a highlighted square. Move one step forward, straight or diagonally. Capture diagonally. Reach the far edge to win.
Learning, measured
03 / SELF-PLAYSame game. Same search budget.
A network that gets better with experience.
Loading measured arena results…
Inspect the evaluation data ↗How does it learn?+
The network predicts a move distribution and who will win. PUCT search explores ahead using those predictions. Two copies play each other; root visit counts become policy targets and the eventual winner becomes the value target. Burn trains both heads from a replay buffer, then the cycle repeats.
This experiment uses a 64-unit network and starts from random weights. The run configuration, source commit and artifact hashes are recorded in experiment provenance. No expert games, heuristic labels, tablebases or solved-game knowledge were used in training.
Learned play is not a proof of optimal play. A separate research result establishes a first-player win for 6×6 Breakthrough. RustZero has not solved the game.
The engine, network, training and search are Rust. Burn Flex performs CPU inference in WASM; a Worker keeps the board responsive. Physical Android performance has not been measured. Source, reproducible Docker commands and validation ↗