← Leo’s experimentsRUST · BURN · WASM
A SMALL SELF-PLAY EXPERIMENT

RustZero↗

A blank slate. A board. A little search.
AlphaZero learns Breakthrough from scratch.

CHAMPION6 × 6 BREAKTHROUGHYOU

Loading Rust engine…

0 plies
No moves yet

Tap a pawn, then a highlighted square. Move one step forward, straight or diagonally. Capture diagonally. Reach the far edge to win.

Learning, measured

03 / SELF-PLAY

Same game. Same search budget.
A network that gets better with experience.

vs heuristic vs random vs heuristic MCTS

Loading measured arena results…

Inspect the evaluation data ↗
How does it learn?+
Network→MCTS→Self-play→Training↺

The network predicts a move distribution and who will win. PUCT search explores ahead using those predictions. Two copies play each other; root visit counts become policy targets and the eventual winner becomes the value target. Burn trains both heads from a replay buffer, then the cycle repeats.

This experiment uses a 64-unit network and starts from random weights. The run configuration, source commit and artifact hashes are recorded in experiment provenance. No expert games, heuristic labels, tablebases or solved-game knowledge were used in training.

Learned play is not a proof of optimal play. A separate research result establishes a first-player win for 6×6 Breakthrough. RustZero has not solved the game.

The engine, network, training and search are Rust. Burn Flex performs CPU inference in WASM; a Worker keeps the board responsive. Physical Android performance has not been measured. Source, reproducible Docker commands and validation ↗