Snake, all the way to a full board.

Watch our frozen neural policy complete all three courses. Slow it down, step through a decision, or jump to the finish.

18 × 18 · green snake · rust apple · grey walls

Watch the board fill

Every move comes from the frozen model’s verified recording. Playback uses the same rules as our Snake game.

Recorded seed 1,900,000 · 324 playable cells

Loading and checking recording…

Recorded demonstration. This page does not run live neural inference or enter human high scores.

What the model learned

The 10,694-parameter network chooses between straight, left and right. It was warm-started from an expert shortcut route, then refined with guided search and reinforcement learning. Its inputs include the board and engineered route geometry.

The measured efficiency gain was already present after expert training. We have not established an extra movement gain attributable to the later RL refinement. Recorded evaluation used unmasked neural argmax, with no teacher override, search or recovery fallback.

What counts as a win?

The snake must fill every playable cell. The three courses contain 324, 316 and 312 cells; walls are excluded. Starting from a three-cell snake, that means collecting 321, 313 or 309 apples. The finish has no remaining food.

The 120 measured wins are validation evidence for these fixed courses, not a guarantee for arbitrary seeds, obstacles or board sizes. The 40.06% movement reduction is from the 60 fresh starts compared on matching initial seeds.

Keep exploring

Model identity and evidence

Version: movement-efficient v3 · training seed 47 · one training lineage. Checkpoint SHA-256:

1a99a8b938fd15cf07071c9c6708b3e650a4726fe1f010629d4fe7cd15dabccd

The three compact recordings reconstruct 43,508 checked states using the website’s TypeScript rules, including each initial state. Downloads are checked by SHA-256 before playback.

Read the complete Snake model card