# Connect Four spatial model with a reduced proof budget

The subsequent policy-teaching round (`reports/2026-10-08-policy-teaching-and-stop.md`, local research reference) also failed its promotion rule. Automatic v1 training is stopped; this checkpoint and controller remain the integration baseline.

The subsequent shared-feature planning trial (`reports/2026-10-08-shared-planning.md`, local research reference) did not pass its promotion gate. This checkpoint and controller remain selected.

The selected local research controller combines the spatial policy/value network with **one million exact-search nodes and 256 neural PUCT simulations**. It passed the reduced-budget acceptance test on fresh paired games: **87.5% score against custom Expert and 70% against Fairy Hard**, compared with 81.5% and 65.5% for the previous five-million-node controller on the same starts. The Hard improvement remains statistically uncertain. This selection recognizes comparable strength with less exact-search work and more reliable tactics.

## Selected artifacts

| Artifact | Identity |
| --- | --- |
| Selected checkpoint (`artifacts/iterate-seed191/selected-candidate.pt`, local research reference) | SHA256 `17d089780bc295fcc91b70ee50efa56659e8495c1004dc78d6692b02a310634f` |
| Controller (`artifacts/iterate-seed191/selected-controller.json`, local research reference) | Exact cap 1,000,000; PUCT 256; partial root proofs enabled |
| Network ONNX (`artifacts/iterate-seed191/export/policy-value.onnx`, local research reference) | 288,371 bytes; SHA256 `2a9887486b37cc21965282ac05ec50772b0deb44a7c264e26896fad27d2d7eb0` |
| Frozen comparison (`artifacts/iterate-seed191/confirmation-comparison.json`, local research reference) | Reduced-budget gate passed; Hard strength-improvement gate did not pass |
| Full experiment report (`reports/2026-10-08-next-iteration.md`, local research reference) | Training, failed variants, controller ablations, uncertainty and reproduction |

Artifacts remain local and ignored by Git. The previous spatial card (`STRONG_MODEL_CARD.md`, local research reference) preserves the earlier checkpoint, controller and historical 72.5% Hard result from a different opening suite. The PPO card (`MODEL_CARD.md`, local research reference) preserves the original actor lineage.

## Architecture and training

The architecture remains the **69,102-parameter** spatial network described in the previous card: 98 visible board and tactical inputs, three residual blocks, seven policy logits and a tanh value for the player to move. Legal masking remains external. Raw inference has no tactical action override. The selected trainable tactical coefficients are approximately 21.446 and 4.856.

The selected model continues the seed 173 spatial champion with training seed **193**, starting from the unchanged parent weights. Seeds 191 and 197 independently shuffle and sample the same replay from that same parent. All three train for 60 epochs; these are three continuation seeds, rather than three independently initialized RL lineages.

Training combines repaired historical replay, **4,096 fresh games**, and certified tactical examples. Fresh games use a depth 10 teacher capped at 100,000 nodes, an exact teacher capped at 250,000 nodes, the frozen parent neural policy, tactical opponents and custom search. The teacher's immediate-win probe can extend its nominal depth by one ply. **1,647 games restart from archived difficult training roots**. The combined dataset contains **191,676 rows from 94,941 canonical positions**, including **161,662 proved value targets**. Unknown values receive zero value-loss weight. Reserved tactical/opening positions and earlier strong-evaluation trajectories are excluded, including reflections and color relabelling.

Immediate wins and genuinely unique safe blocks take precedence in policy teaching, including positions with a proved eventual loss. This repaired 1,682 historical policy rows. Training uses masked cross entropy, proved-value MSE and a tactical logit margin of 2. Every epoch presents all rows plus tactical replay equal to half the dataset. AdamW uses learning rate 0.00015 for ordinary parameters and 0.003 for the two tactical coefficients, weight decay 0.0001, batch 512 and gradient clipping 2. Training ran on local Apple MPS. This phase is supervised search distillation with an archived-root curriculum.

## Fresh results

Each opponent played **100 games**, swapping seats on 50 unique 4–12-ply starts. Score gives draws half credit. Both controllers use identical openings and paired seeds 780191–780240. The actual bundled Fairy engines retain their app settings and skill randomness.

| Opponent | Previous controller | Selected controller | Selected 95% interval |
| --- | --- | --- | --- |
| Custom Level 2 | 83.5% | 89% | 83–94% |
| Custom Level 3 Expert | 81.5% | 87.5% | 81.5–93% |
| Window six-ply search | 74% | 79.5% | 72.5–86.5% |
| Fairy Normal | 84.5% | 89% | 83–94% |
| Fairy Hard | 65.5% | 70% | 63.5–76.5% |

Paired score changes against Level 2 and Expert are +5.5 points [1.5,10] and +6 points [1,11.5]. Against Hard, the change is **+4.5 points [−3,12]**, so a strength improvement is unproved. The fixed acceptance rule required the lower paired interval to exceed −5 points against Expert, window 6 and Hard at the smaller budget, at least 99% raw tactical accuracy, and no mean regression greater than five points against any reference. It passed.

On **800 additional fresh tactical cases**, raw inference takes 400/400 immediate wins and 400/400 safe blocks; the parent takes 400/400 and 374/400. Raw gameplay remains weak: in a separate 60-game development suite, the selected raw network scores 29.2% against Expert, 20% against window 6 and 19.2% against Hard. Serving ONNX argmax alone does not reproduce the hybrid scores.

## Search work and verification

Partial exact proofs constrain root actions when a query cannot establish the whole position. Proved losing actions are excluded while unresolved actions remain available. The remaining search behavior follows the previous card: bounded depth 12 terminal-win search with a 150,000-node cap, then neural PUCT with exploration 1.5, prior temperature 1 and a 5% support floor. Proof transpositions reset per game; only public move history is used.

The confirmation used **607,225,254 exact nodes**, versus 1,681,607,027 for the parent: **63.9% less exact-search work** across the same 500 games. Exact proofs handle 5,920/6,361 decisions, about 93.1%; search still contributes most of the strength. On 40 unique cold-cache public positions, p95 latency falls from **274.6 to 100.5 ms**, while median latency rises from 64.5 to 88.5 ms. The smaller cap reduces long waits but sends more positions to MCTS.

**76 tests pass**. All **6,716 completed games and 196,541 transitions** match both live TypeScript game engines. Training audits found zero reserved-position overlap and validated legal targets, tactical labels, exact value receipts and reflections. ONNX Runtime agrees on 1,000/1,000 actions, with individual and dynamic-batch errors below 1e−4. The export contains the network; native proof search is separate.

The model remains a local research selection. Final seeds 900000–999999 remain sealed, browser/mobile worker integration is pending, and the three continuations share a parent and training data. The local controller requires prepared native teachers and roughly 80 MiB of exact-solver cache. No hosted inference or recurring service was added.
