# Snake model card

The current neural Snake model fills all three courses and uses about **40% fewer movements** than the preserved cycle model. It won **120/120 full-board benchmark starts** on 7 October 2026. It learned a shortcut expert before receiving reinforcement learning and search updates; the measured movement improvement was already present after expert training.

This is a local research model for the game's normal starting state. Browser integration and qualification across independent training seeds remain pending.

## Model identity

| Property | Current model |
| --- | --- |
| Version | Movement-efficient v3; observation `snake-efficiency-spatial-v3` |
| Game rules | `snake-fill-board-v2`, implemented in the reference game (`../../../lib/snake.ts`, local research reference) and Python simulator (`engine.py`, local research reference) |
| Method | Shortcut expert warm-start, followed by stochastic PUCT policy/value iteration |
| Training seed | 47; one training lineage |
| Selected checkpoint | selected.pt (`artifacts/efficiency/paper-seed47/selected.pt`, local research reference), an unchanged copy of `round-001.pt` |
| Portable graph | actor-value.onnx (`artifacts/efficiency/paper-seed47/export/actor-value.onnx`, local research reference), 47,587 bytes |
| Encoder contract | schema.json (`artifacts/efficiency/paper-seed47/export/schema.json`, local research reference) |
| Verification receipt | manifest.json (`artifacts/efficiency/paper-seed47/export/manifest.json`, local research reference) |
| Detailed evidence | Measured report (`reports/2026-10-07-efficiency-v3.md`, local research reference) and all numerical results (`reports/2026-10-07-efficiency-v3.json`, local research reference) |

SHA-256 identifiers:

```text
selected.pt:      1a99a8b938fd15cf07071c9c6708b3e650a4726fe1f010629d4fe7cd15dabccd
actor-value.onnx: 06f8aa29d026bbf21e7a64862942fb17e042aac0e3fde4ce22ec22c9b44043cb
schema.json:      73e985d9ffb866a25caf93151b8dc9eb89ce42300dd3bae599a6b041e524bcb7
```

Large checkpoints, exports, traces and virtual environments live in ignored local `artifacts/` directories. A source checkout alone does not contain these weights. Keep the measured run and its frozen `source/` copies intact; use a new directory for another run.

## Game and objective

The board is 18×18. The snake starts with three cells, facing right, at `(5,3), (4,3), (3,3)`. Completion requires occupying every playable cell and exhausting food; collecting a small target number of apples does not complete the current game.

| Course | Playable cells | Fresh apples required |
| --- | ---: | ---: |
| Open garden | 324 | 321 |
| The long path | 316 | 313 |
| Twin hedges | 312 | 309 |

Actions are relative to the current heading: **0 = straight, 1 = left, 2 = right**. In measured inference the network chooses the largest of its three raw logits. All actions remain available, including unsafe ones. There is no teacher call, search, action mask, replacement action, continue or fallback.

Checkpoint selection ranks full-board wins first, then mean movements among wins. Evaluation reports collision losses, 100,000-movement cutoffs and 2,000-movement inactivity cutoffs separately. Difficulty changes timing, not movement rules or the completion target; the benchmark uses difficulty 1.

## Inputs and architecture

The encoder (`efficient_policy.py`, local research reference) produces three float32 inputs. Neither the game seed nor hidden random-generator state enters the model.

| Input | Shape, excluding batch | Contents |
| --- | --- | --- |
| Spatial board | `9×18×18` | `body_up`, `body_right`, `body_down`, `body_left`, `head`, `tail`, `food`, `walls`, `ordered_body`, in this order |
| Extra state | `8` | Four heading bits, three course bits, and body length / course capacity |
| Candidate geometry | `3×18` | Visible geometric features for straight, left and right, derived from the current body, food, tail and a certified course cycle |

Non-head body directions point toward the preceding segment; the head uses its current heading. Ordered segment `i` stores `(body_length-i)/324`. The actor's spatial order is the top-level `board.planes` order in the export schema; the nested geometry helper's `state_planes` describe a separate helper representation.

The 18 geometry features, in order, are:

```text
legal_now, ordered_body_now, preserves_order, food_not_skipped,
tail_not_overtaken, growth_room, teacher_admissible, eats_now,
advance_fraction, food_forward_fraction, tail_forward_fraction,
effective_tail_forward_fraction, next_food_forward_fraction,
next_tail_forward_fraction, free_forward_after_fraction,
growth_reserve_fraction, free_cells_fraction, food_manhattan_delta
```

The first 17 lie in `[0,1]`; the last lies in `[-1,1]`. These engineered inputs give the model substantial prior knowledge of safe cycle shortcuts. `teacher_admissible` is an input flag, not an inference mask. Geometry is computed outside the ONNX graph.

The network has **10,694 learned parameters**:

- Spatial encoder: two 3×3 convolution layers, `9→8→8` channels, ReLU activations, average pooling to 3×3, then `72→24` with ReLU. Appending the eight extra values gives a 32-value context.
- Actor: each action's 18 features receive a learned linear score plus `0.1 × tanh(residual)`. The shared residual network uses `50→64→32→1` layers with ReLU hidden activations, taking action features and context.
- Value head: `32→64→3`, predicting a win logit, discounted return divided by 5, and log remaining movements divided by `log(100001)`. Only actor logits choose evaluation actions.

The selected weights have a checked admissibility dominance margin of **0.590878** under the declared feature bounds. An admissible candidate therefore outranks unflagged candidates when one exists. This bound does not prove universal solvability or recovery from arbitrary saved states. The win head is trained mainly on certified winning trajectories; its off-domain probabilities are not calibrated.

## Training and checkpoint selection

The configuration (`configs/efficiency-paper.json`, local research reference) and training implementation (`efficient_train.py`, local research reference) record this recipe:

1. Generate 18 full games with a shortcut expert (`shortcuts.py`, local research reference), using seeds `504700..504717`. The teacher preserves cyclic body order and growth space and stops taking shortcuts above half board fill. Its 264,460 movements yield 73,744 sampled teaching records, balanced by course, body phase and steering action.
2. Train for 35 warm-start epochs with Adam at `0.003`, batches of 256 and gradient clipping at 5. The objective combines policy cross-entropy, win binary cross-entropy, return SmoothL1, `5 ×` movement SmoothL1 and `3 ×` admissibility-margin regularization.
3. Generate nine guided policy-improvement games, seeds `554700..554708`. Every three games, train five epochs on the latest 40,000 records at learning rate `0.0005`. Adam is recreated for each training phase.

PUCT search (`efficient_search.py`, local research reference), which balances predicted value and exploration, runs eight simulations at movement multiples of 128 when there are multiple admissible actions. It uses exploration coefficient `0.5`, depth 24 movement ticks and forced sequences up to 32 ticks. Every physical movement counts in both the budget and discount exponent. Future apples are independently sampled uniformly from free cells; search discards the real hidden game seed.

Training rollouts and search restrict choices to the certified cyclic-order subset. **The benchmark actor performs raw, unfiltered inference.** Rewards are `1/(capacity-3)` per apple, `+5` for completion, `−5` for collision and `−0.00001` per movement, with discount `0.9999`. These soft training rewards do not guarantee that winning always outranks speed; explicit checkpoint selection supplies that priority.

| Work | Whole experiment | Selected checkpoint's lineage |
| --- | ---: | ---: |
| Expert movements | 264,460 | 264,460 |
| Real policy-improvement movements | 133,112 in 9 games | 43,435 in 3 games |
| Simulated search movements | 5,313 | 1,967 |
| Refinement rounds | 3 | First round |

The measured experiment took 416.33 seconds on one CPU/PyTorch thread. Its reported training counter totals 402,885 movements. Twelve probes add 177,428 evaluation movements, giving **580,313 total movements** across these categories, below the 1,800,000 cap. Subsequent guardrail repairs include probes in future budget accounting; frozen sources preserve the measured implementation.

Warm-start and all three refinement checkpoints each won 15/15 selection starts, with identical mean movements of 14,565.13. These starts use `1860000 + course×10000 + i`, `i=0..4`. Ties prefer a checkpoint after refinement, then the earliest round, selecting `round-001.pt`. Its weights changed by L2 distance 1.2783 from warm-start. The measured speed-up came from learning the shortcut expert; additional movement improvement from RL/search was not observed on selection starts.

## Evaluation results

Fresh validation uses `1900000 + course×10000 + i`, `i=0..19`. The comparator is the preserved v2 policy's validated cycle route, replayed on matching starting seeds. Different paths produce different later apple locations; matching starting seeds does not imply identical apple sequences.

| Fresh course | Cycle mean movements | Current mean movements | Reduction | Full-board wins |
| --- | ---: | ---: | ---: | ---: |
| Open garden | 25,815.30 | 15,436.50 | 40.20% | 20/20 |
| The long path | 24,784.15 | 14,788.95 | 40.33% | 20/20 |
| Twin hedges | 23,941.55 | 14,455.25 | 39.62% | 20/20 |

| Suite | Wins | Previous → current mean movements | Reduction | Current median / worst winning movements |
| --- | ---: | ---: | ---: | ---: |
| Fresh 60 starts | 60/60 | 24,847.00 → 14,893.57 | 40.06% | 14,786.5 / 16,464 |
| Existing 60 starts, base seed 1,800,000 | 60/60 | 24,814.00 → 14,754.60 | 40.54% | 14,752 / 16,276 |

Neither suite contains deaths, inactivity cutoffs, movement cutoffs or assisted runs. The fresh reduction's paired, stratified bootstrap 95% interval is **39.31%–40.79%**. It describes starting-seed variation for this fixed model, not variation across independently trained models. These are validation suites; they do not establish an optimal path or an improvement over the shortcut teacher.

## Correctness and export evidence

- **64 Python tests passed**, including ten efficiency and nine shortcut regressions. Both TypeScript parity bridges passed ESLint. Game rules were unchanged in this efficiency experiment.
- The first fresh start on each course takes **15,003 / 14,659 / 13,843 movements**. All **43,508 initial/resulting states** match the live TypeScript state hash chains, ending with full bodies and null food.
- ONNX matches the authoritative trained policy on **1,000 unmasked actions**, spanning every course and body lengths 3–323. Maximum errors are `2.38e−6` for logits and `3.81e−6` for values, below the `1e−4` tolerance.

Training uses Python 3.12.13 and PyTorch 2.14.1. Export reuses Connect Four's existing PyTorch 2.8.0, ONNX 1.19.0 and ONNX Runtime 1.23.0 environment. Export parity compares against saved outputs from the actual Snake training runtime, accounting for that version difference. Dependency pins are in requirements.lock (`requirements.lock`, local research reference).

Browser/mobile geometry encoding, inference latency and production integration remain unverified. Qualification also needs independent training seeds and the sealed final-test seeds at or above 2,000,000. The current results apply to the three courses' aligned factory starts; arbitrary saved-state recovery is outside the measured scope.

## Reproduction

Run from `experiments/rl/snake`. See setup and the complete training, selection and export sequence (`README.md#reproduce-efficiency-work`, local research reference). The scripts do not read the JSON configuration automatically; the README supplies the flags matching that recipe. Always use fresh output paths, including for evaluations that can replace result files.

Inspect the existing winner on the same fresh validation suite without training or replacing it:

```sh
.venv/bin/python efficient_eval.py --run artifacts/efficiency/paper-seed47 --models selected.pt --episodes 20 --seed-start 1900000 --out artifacts/efficiency/card-validation-fresh-v3 --traces
```

The tests run with:

```sh
.venv/bin/python -m pytest tests -q
```

Reproduction may produce different weights or timings; it does not substitute for an independent-seed qualification run. Keep `selected.pt`, prior candidates and the export receipts immutable.

## Preserved model history

| Version | Method and measured outcome | Scope and evidence |
| --- | --- | --- |
| Legacy v1 | DQN pilots and repair; later 194/200 wins at a 12-apple goal | Former 12/18/24-apple goals, not full-board completion. Pilot (`reports/2026-10-07-pilot-v1.md`, local research reference), selected demo (`reports/2026-10-07-long-path-demo-v2.md`, local research reference), reliability (`reports/2026-10-07-legacy-reliability-v1.md`, local research reference) |
| Full-board v2, seed 31 | Expert cycle warm-start + 200,000 DQN steps; 60/60 full-board wins | Head/heading/course selector with 331 inputs, two 128-unit ReLU layers and three Q values. Food/body-independent route cannot learn shortcuts. Report (`reports/2026-10-07-fullboard-hybrid-v2.md`, local research reference) |
| Movement-efficient v3, seed 47 | Spatial actor/value model, shortcut warm-start + search/RL; 120/120 wins and about 40% fewer movements | Current model documented above. Report (`reports/2026-10-07-efficiency-v3.md`, local research reference) |

The v2 checkpoint is `artifacts/fullboard/hybrid-stable-seed31/model.zip`, SHA-256 `8a529fd77db142377533f2eff5c356cc11f7b8e1f8caabe4cc062ecbb60c0077`. Its actions matched all 952 certified route keys; its three complete demonstrations matched 73,381 TypeScript states. An earlier weak refinement regressed route actions and remains preserved as a failed candidate. Legacy sources are isolated under `legacy_v1/`; configs without a rules version resolve to legacy rules.

The [current private video](https://youtu.be/30RMD21jI2I) shows all three first fresh-validation wins at labelled 100× time-lapse. The [preserved v2 private video](https://youtu.be/XwMVTL6S6fE) shows the earlier cycle model. Frames sample complete verified movement traces; these are accelerated replays of neural actions.

## Research references

[AlphaSnake](https://arxiv.org/abs/2211.09622) informed spatial observations, policy/value search and movement-aware discounting. [John Tapsell's shortcut algorithm](https://johnflux.com/2015/05/02/nokia-6110-part-3-algorithms/) informed cyclic-order and growth-space checks. Our board, expert supervision and training restrictions differ from the paper's experiment; this model is an adaptation, not a reproduction of its reported results.
