# Connect Four shared feature planning trial

The next iteration improved value calibration but **did not establish stronger gameplay**, so the selected champion remains unchanged. The frozen candidate scores 71% against Hard at half the exact-search cap, versus 70.5% for the champion at its normal cap. The paired interval is too wide to pass the fixed noninferiority rule, and the unchanged champion scores 72.5% at the same smaller cap.

The candidate model card (`../PLANNING_MODEL_CARD.md`, local research reference) documents the exact checkpoint, architecture, training and runtime contract. The machine report (`2026-10-08-shared-planning.json`, local research reference) preserves the measured ledger. These are local supervised search-distillation experiments; no additional PPO phase or production deployment occurred.

## Larger neural MCTS comparison

The earlier 20-game comparison suggested a 45% → 60% Hard improvement. This iteration first tested both frozen networks on 100 games per opponent, using 50 fresh unique starts with swapped seats, the same 256-simulation PUCT settings, and no exact solver.

| Opponent | Older spatial network | Current selected network | Paired change and 95% interval |
| --- | --- | --- | --- |
| Custom Expert | 64.5% | 76.5% | +12 points [2.5,21] |
| Window six-ply search | 46% | 60.5% | +14.5 points [5.5,23.5] |
| Actual Fairy Hard | 34.5% | 36% | +1.5 points [−8.5,12] |

The lower-level gains hold up, but the large Hard gain does not. The smaller result was development evidence rather than a general reliability estimate. The actual bundled engines retain their app settings and skill randomness. Game seeds are 790211–790260, with opening seed 791211; final seeds remain sealed.

## Shared feature training and overfitting

This trial adapts [Expert Iteration](https://arxiv.org/abs/1705.08439) by teaching policy/value networks from stronger offline planning, and the archived-position approach of [Go-Exploit](https://arxiv.org/abs/2302.12359) by reusing difficult public training roots. The earlier frozen-feature value update failed; this time the shared trunk learns alongside both heads.

The first collection supplies 1,024 early training positions, mixing archived unresolved roots with fresh 6–16-ply prefixes. An independent 400-position value-validation set is excluded from all replay, including mirrored and color-equivalent boards; it contains 388 proved outcomes. Its positions are also absent from parent training. Each exact query has a 20-million-node cap. Heuristic scores never become value labels.

Three continuations train for 60 epochs with heavy early-position repetition. Their training losses improve, but validation error rises. For example, seed 223 goes from 0.508 at epoch 10 to 0.594 at epoch 60, versus 0.516 for the parent. All retain 800/800 tactical accuracy, but a controller scout rejects the most promising early checkpoint. The recipe is preserved as a failed attempt.

The next collection adds 8,192 distinct positions unseen by the parent, proving 7,945. Four bounded collection processes complete this in about 218 seconds for the slowest shard. The expanded replay contains 58,289 rows and 49,438 proved values. Early rows receive one extra presentation per epoch instead of eight. Strategic replay distils the current policy, certified tactics retain their labels, and policy KL limits drift. Unproved roots supply 512-simulation visit targets with any known action proofs; unproved values remain masked.

All three broader continuations improve value-validation error to approximately 0.43 and pass the 800-case tactical development suite. Increasing epochs to 30 does not improve validation error over epoch 10. Seed 257 at epoch 10 is frozen for confirmation based on those checks and the paired controller scouts. Its error is 0.434 versus 0.516 for the parent, about 15.9% lower, but outcome classification changes only 71.9% →72.2%. Better confidence did not yield a clear improvement in classification.

## Frozen gameplay confirmation

The preregistered targets were 80% against Hard at one million exact nodes with a positive paired lower bound, or comparable performance at 500,000 nodes. The latter required lower paired bounds above −5 points against Expert, window six-ply search and Hard, perfect raw tactics, and no reference mean regression greater than five points.

The candidate uses 500,000 exact nodes, 256 PUCT simulations and partial root proofs. The champion uses one million nodes with the same planning settings. Both play 100 games per opponent on identical 50 unique 4–12-ply starts, swapping seats, with seeds 840211–840260 and opening seed 841211. A separate unchanged-parent comparison uses the candidate's 500,000-node cap.

| Opponent | Champion 1m | Candidate 500k | Parent 500k |
| --- | --- | --- | --- |
| Custom Level 2 | 87.5% | 88% | — |
| Custom Expert | 87% | 86% | 86% |
| Window six-ply search | 74.5% | 75% | 76.5% |
| Fairy Normal | 92% | 88.5% | — |
| Fairy Hard | 70.5% | 71% | 72.5% |

The Hard change versus the normal-budget champion is +0.5 points [−7,8]. Expert is −1 point [−5,3], and window six-ply search is +0.5 points [−6,7]. All three miss the required strict lower bound. The mean/tactical guards pass, including 400/400 immediate wins and 400/400 safe blocks on new fixtures, but the candidate is not promoted. Against the parent at equal budgets, score changes are 0, −1.5 and −1.5 points; every interval includes zero.

The new network with MCTS alone scores 42.5% against Hard, versus 36% for the selected network, on the earlier development bank. The paired gain is 6.5 points [−3.5,16.5], still inconclusive. This bank has already been used for development and is not presented as untouched final qualification. No controller reaches the 80% Hard target in this trial.

## Verification and artifacts

All 83 tests pass, including unknown-value gradient masking, shared-feature updates, unchanged tactical coefficients, strict runtime budgets and frozen acceptance rules. All 2,480 measured complete games match both live TypeScript engines on 82,052 transitions. The 9,616 early-position prefixes add 106,991 matching transitions. Training audits match proof receipts and mirrored labels, verify input encoding and legal policies, and find zero reserved-position overlap. The expanded 8,192 positions and value-validation cases are absent from parent training.

The candidate ONNX matches 1,000 PyTorch actions and dynamic batch 17; maximum logit/value/batch errors remain below 1e−4. It exports the network rather than the native search controller. All six continuation seeds, intermediate checkpoints, teaching targets, query receipts, game traces, opponent commands, source snapshots and decision files remain under `artifacts/planning-seed211/`. The selected champion's bytes remain unchanged. The trial shares one initial training lineage, and final seeds 900000–999999 remain unused.

Run the frozen candidate locally from this directory:

```sh
.venv/bin/python -m pytest -q
.venv/bin/python strong_play.py --model artifacts/planning-seed211/candidate.pt --controller artifacts/planning-seed211/candidate-controller.json --moves 0,6,1,6,2,5
```

The next experiment should focus on whether stronger search policy targets improve action selection at a fixed budget. Value MSE alone is an inadequate promotion objective: this trial improved it without establishing a strength gain. The candidate remains useful as a calibration control; the selected model stays documented in the current card (`../ITERATION_MODEL_CARD.md`, local research reference).
