Fill the same board with fewer movements
The selected Snake policy used 40.06% fewer moves than the cycle baseline on fresh measured starts.
· Leaf Arcade experiment notes
A win comes before efficiency
A short run that crashes is not an efficient solution. Selection first maximizes full-board wins and then compares movement counts among winners. Every comparison uses the same course and starting seed.
The measured improvement
Across the recorded fresh comparison, mean movements fell from 24,847 for the cycle baseline to 14,893.57 for the selected neural policy: a 40.06% aggregate reduction. Across its measured local validation sets the selected model finished 120 out of 120 starts.
The public recordings are the fixed first seed of each fresh course stream, not the shortest run found after searching for a favourable demonstration.
Where the improvement came from
Expert shortcut behaviour supplied the warm-start. Guided search and reinforcement-learning refinement changed the model, but the warm-start and refined checkpoints tied on the selection suite. We have not established an additional movement improvement caused by RL after that expert initialization.
At recorded evaluation time, the frozen neural policy chose the largest of three logits without a teacher override, action filter, search, continue or fallback. Its input still contains engineered route geometry. That prior knowledge is part of the method.
Ideas to read next
- AlphaSnake: Policy Iteration on a Nondeterministic NP-hard Markov Decision Process
- John Tapsell: Snake cycle and shortcut algorithm
These sources informed the method. Our small experiments do not reproduce their full research results.