# Movement efficiency experiment — 7 October 2026

User objective: preserve full-board wins while reducing movement count. The existing `hybrid-stable-seed31` model and all prior videos/checkpoints remain preserved. This is a new local experiment, not a production gameplay change.

## Research applied

[AlphaSnake](https://arxiv.org/abs/2211.09622) motivates the win-first objective, spatial body/head/tail/apple observations, a policy/value network, PUCT visit-count teaching, stochastic future-food branches and movement-aware discounting of forced-action sequences. Its reported benchmark uses a 10×10 board and a 1,200-step limit. Our three 18×18 courses, expert start, training restrictions, budget and evaluation protocol differ; we do not claim to reproduce its reported performance.

[John Tapsell's original shortcut algorithm](https://johnflux.com/2015/05/02/nokia-6110-part-3-algorithms/) informs the training expert's cyclic body-order and growth-space checks. The expert is measured separately from learned inference.

## Implemented adaptation

The actor/value network sees nine full-size planes: four body directions toward the head, head, tail, current apple, walls and ordered body. Eight additional values describe heading, course and board fill. Each of the three actions also has 18 public geometric features derived from visible state and a fixed course cycle. These include current legality, cyclic order, growth space, current-food distance and advance. Hidden RNG state and future apples are excluded.

The actor combines learned linear geometric weights with a bounded nonlinear spatial residual. A training regularizer encourages a positive learned-logit dominance margin for admissible geometry. It changes parameters through ordinary gradient updates; it never projects weights or replaces, filters or masks inference actions. The resulting bound is checked after training and applies only to the declared feature ranges and existence of an admissible candidate. It is not a universal game-solvability proof.

Warm-start uses separately tested shortcut expert trajectories. Policy improvement uses real game rollouts, search visit distributions and discounted-return/remaining-movement targets. Its training search explores the explicitly declared cyclic-order subset, with independent uniformly sampled future food. Single-choice sequences retain every simulated movement and the correct discount exponent. Raw neural evaluation contains no search, teacher, controller, action mask or fallback.

Rewards are normalized apple progress, +5 for actual full-board completion, −5 for collision, and −0.00001 per movement. Discount is 0.9999. These are training surrogates. Candidate selection explicitly compares full-board wins first, then mean winning moves on the same fixed starts; failed short games cannot improve the movement statistic. All failures and cutoffs remain reported.

## Budget and protocol

One CPU/PyTorch thread, existing locked dependencies. Fresh output directory; existing artifacts cannot be overwritten. Configuration is efficiency-paper.json (`../configs/efficiency-paper.json`, local research reference). Initial bounded phase: 18 expert games, 35 supervised epochs, 9 policy-improvement games, 8 search simulations every 128 movement ticks, up to 1,800,000 total expert/real/search ticks and 1,000 wall seconds.

Training seeds stay below 600,000. Probe starts use 1,760,000+course×10,000; selection starts use 1,860,000+course×10,000; independent fresh validation uses 1,900,000+course×10,000. Final-test seeds at/above 2,000,000 remain reserved. Pair identical starting seeds against the frozen current cycle winner. Report all-start win rate, average/median/worst winning moves and per-course reduction. Different paths naturally cause different subsequent food positions; comparisons do not assume identical apple locations.

## Measured neural results

**The selected raw neural model won 120/120 benchmark starts and reduced movement count by about 40%.** It met the proposed 25% reduction target while retaining every full-board win on the existing 60-start suite, then passed an independent fresh 60-start suite. Every course has 20 starts per suite, with no losses, cutoffs, assistance or replacement actions.

| Fresh course | Previous policy mean moves | New neural mean moves | Paired reduction | Full-board wins |
| --- | ---: | ---: | ---: | ---: |
| Open garden |25,815.30 |15,436.50 |40.20% |20/20 |
| The long path |24,784.15 |14,788.95 |40.33% |20/20 |
| Twin hedges |23,941.55 |14,455.25 |39.62% |20/20 |

Fresh aggregate: 24,847.00→14,893.57 mean moves, **40.0589% reduction**. Existing suite: 24,814→14,754.6, **40.5392% reduction**. Median and worst winning moves, every individual result, all candidate outcomes and stratified paired bootstrap intervals are in the machine-readable report (`2026-10-07-efficiency-v3.json`, local research reference). Intervals describe starting-seed variability for this fixed model; independent training-seed variation remains unmeasured.

Four immutable candidates (warm-start plus three refinement rounds) each won 15/15 selection starts with identical movement statistics: 14,565.133 mean moves and 41.0616% reduction. Explicit tie-breaking selects the earliest post-RL checkpoint. The **warm-start already supplied the observed speed-up** on these starts; policy/value iteration retained it. We do not attribute additional movement gains to RL/search or claim an advantage over the shortcut expert.

Selected source: `artifacts/efficiency/paper-seed47/round-001.pt`, copied unchanged to `selected.pt`, SHA-256: `1a99a8b938fd15cf07071c9c6708b3e650a4726fe1f010629d4fe7cd15dabccd`. It has 10,694 learned parameters and a positive raw-logit dominance margin 0.590878 under the declared geometric bounds. Its lineage includes 264,460 expert movements, 43,435 real guided-policy movements and 1,967 search movements, followed by five policy/value training epochs. Parameters changed L2=1.2783 after warm-start.

The whole bounded experiment completed 402,885 movements: 264,460 expert, 133,112 real policy-improvement and 5,313 search movements, in 416.33 seconds. All nine policy-improvement games won. Search visited 664 nodes, sampled 73 new-apple outcomes and counted 4,102 forced movements. Later candidates and their reports remain preserved; the final experimental network changed L2=3.1111, distinct from the selected lineage.

## Verification and portable export

The complete first fresh-start traces have 15,003/14,659/13,843 movements. Every initial/resulting state matches the live TypeScript hash chain: **43,508 states, zero mismatches**, ending at 324/316/312 cells with null food. Raw unmasked actor logits and every action are recorded.

The final Python suite passed **64 tests**, including all ten focused efficiency regressions and nine shortcut checks. Shortcut checks cover tail vacancy/growth, food/tail order, seed exclusion and real full-board rollouts. Both TypeScript parity bridges passed ESLint. Existing game rules were unchanged, so prior production gameplay checks were reused.

ONNX actor/value graph: `artifacts/efficiency/paper-seed47/export/actor-value.onnx`, 47,587 bytes, SHA-256: `06f8aa29d026bbf21e7a64862942fb17e042aac0e3fde4ce22ec22c9b44043cb`. It matched the authoritative trained policy on **1,000 unmasked actions**, with maximum logit error 2.38e−6 and value error 3.81e−6. Fixtures span body lengths 3–323 and all courses. Export reused the existing Connect Four environment (torch 2.8.0/ONNX 1.19.0/Runtime 1.23.0), and was checked against independently saved outputs from the actual Snake training runtime (torch 2.14.1). This version difference is recorded; no new dependencies were installed.

The schema records the complete spatial and geometric encoder contract. Its geometry calculations remain outside ONNX. Browser/mobile encoder parity, latency and deployment are still outstanding. This is one hybrid training lineage evaluated from aligned factory starts, with engineered cycle knowledge; arbitrary saved-state recovery and a calibrated off-domain win probability are not established. Final qualification seeds remain sealed.

## Reproduce without overwriting evidence

Use a NEW output directory. The existing selected/checkpoint directories are preserved.

```sh
.venv/bin/python efficient_train.py --out artifacts/efficiency/reproduction-seed47 --warm-games 18 --warm-epochs 35 --rl-games 9 --wall-seconds 1000 --max-training-ticks 1800000 --search-every 128 --search-simulations 8
.venv/bin/python efficient_eval.py --run artifacts/efficiency/reproduction-seed47 --episodes 5 --seed-start 1860000 --out artifacts/efficiency/reproduction-seed47/selection
```

Reuse candidate result JSONs with `efficient_select.py` to apply the exact-tie rule, then run the frozen selection on fresh 1900000 streams and the existing 1800000 benchmark. Export uses verified trace files with `efficient_export.py --make-fixtures ...`, followed by`--onnx` in the existing export environment.


Post-run review repaired future-run guardrails: neural probes now participate in the movement/wall budget, and evaluation no longer offers an alternate selection path that could overwrite the saved winner. The completed run stayed within its cap before these repairs. Its 12 recorded probe games add 177,428 evaluation movements; total expert/real/search/probe work was 580,313 movements, still below 1,800,000. Frozen sources for the measured run are preserved under its `source/` directory.

Verified video upload: [Private YouTube — faster neural Snake](https://youtu.be/30RMD21jI2I). It shows the three fixed first fresh-validation starts, 43,508 state matches and all playable cells filled, at labelled 100× time-lapse. The previous v2 video remains separate.
