# Connect Four policy teaching and v1 training stop

The bounded policy-teaching round did not establish an accepted strength gain. Its frozen candidate scores **77% against Fairy Hard**, versus **75% for the selected champion**, at identical one-million-node and 256-simulation budgets. The paired change is **+2 points [−4.5,8.5]**. The current champion remains selected, and automatic v1 training stops after three consecutive hypotheses fail their promotion rules.

The ML Lab product plan (`../../../../docs/ml-lab-product-plan.md`, local research reference) separates model quality from browser readiness. Existing models are sufficient for an honestly labelled educational preview. Readable experiment notes, verified demonstrations and runtime integration now offer more product value than another unspecified hill climb. An 80% Hard score remains an optional research target rather than a launch prerequisite.

## Bounded experiment

This round tested one hypothesis with two policy temperatures, one small scout and one frozen fresh confirmation. It reused 768 eligible strategic training positions from the earlier teaching archive, excluding previous evaluation trajectories and new reserved fixtures. No additional large training dataset was generated.

The teacher uses nominal depth 14 with a two-million-node cap, plus the existing immediate-win probe. It completed depth 14 on 543 positions, depth 12 on 223 and depth 10 on 2. Additional child-position exact queries, capped at 500,000 nodes each, classified every legal action on 530 positions. Child values are negated to the parent player's perspective. Unresolved child actions remain unlabelled; only guaranteed optimal actions enter policy targets.

Unlike earlier one-action targets, the new targets distribute probability across admissible actions using deeper search scores. The two variants use temperatures 80 and 240. Terminal-win proofs take precedence over heuristic soft targets. Both train only the policy projection and head for 30 epochs, starting from the selected checkpoint. Shared features, tactical coefficients and the complete value function remain unchanged. Strategic replay and certified tactical examples limit drift.

Collection took 66.9 seconds locally. Both variants passed 800/800 development tactics. The soft variant matched the champion against custom Expert and window 6 in the 24-game-per-reference scout, while Hard score rose from 58.3% to 62.5%. It was frozen before the separate confirmation; no further variants or confirmation retries were used.

## Fresh comparison

Each controller played 100 games per opponent on the same 50 unique 4–12-ply boards with swapped seats, game seeds 860271–860320 and opening seed 861271. Draws receive half credit. The actual bundled engines retain their app difficulty settings and skill randomness.

| Opponent | Champion | Candidate | Paired change and 95% interval |
| --- | --- | --- | --- |
| Custom Level 2 | 87.5% | 88.5% | +1 point [−3.5,5.5] |
| Custom Expert | 88.5% | 88% | −0.5 points [−5.5,4] |
| Window six-ply search | 80.5% | 84.5% | +4 points [−1,9] |
| Fairy Normal | 91.5% | 94.5% | +3 points [−1.5,7.5] |
| Fairy Hard | 75% | 77% | +2 points [−4.5,8.5] |

The preregistered strength gate required at least three points against Hard with a positive paired lower bound, perfect fresh raw tactics and no mean reference regression above five points. The candidate passes the tactical and mean guards, including 400/400 immediate wins and 400/400 safe blocks, but fails the Hard requirement. No opponent interval establishes a clear gain. The promising candidate remains available as a research artifact rather than replacing the champion.

## Artifacts and verification

| Artifact | Identity |
| --- | --- |
| Frozen candidate (`../artifacts/policy-seed271/candidate.pt`, local research reference) | SHA256 `bc9d3256573931faef99963c11a3cc07295de54fa1faf0f36544d02585bb9c09` |
| Candidate controller (`../artifacts/policy-seed271/candidate-controller.json`, local research reference) | Exact cap 1m; PUCT 256; partial root proofs |
| Candidate ONNX (`../artifacts/policy-seed271/candidate-export/policy-value.onnx`, local research reference) | 288,371 bytes; SHA256 `5b8a37ec943898e2e71676710f972ac77cd129a97c22ba1f0a33caa31fd648f5` |
| Decision (`../artifacts/policy-seed271/decision.json`, local research reference) | No promotion; automatic v1 training stopped |
| Machine report (`2026-10-08-policy-teaching-and-stop.json`, local research reference) | All outcomes, teacher coverage and provenance |

All 86 tests pass. All 1,216 completed games match both live TypeScript engines on 41,212 transitions; the 768 teaching prefixes add 8,436 matching transitions. Audits validate action proofs, child-value signs, legal targets, horizontal reflections, zero reserved-position overlap and identical non-policy parameters. ONNX matches 1,000 PyTorch actions; individual and batch 17 errors stay below 1e−4. Only the network is exported, not the native search controller.

The selected checkpoint and its settings remain unchanged. Final seeds 900000–999999 remain sealed. The local experiment uses two controlled variants sharing one parent, not independent initial-training qualification. No browser gameplay, hosting, billing or public publication changed.

## Next product step

Freeze the accepted Snake and Connect Four models for v1. Build the ML Lab journal and verified Snake replay first, then qualify live browser encoders, workers, cancellation and response times. Keep neural-only, neural-plus-lookahead and native-solver results separately labelled. Supported seed, course, speed, stepping and bounded thinking controls provide useful player experimentation without hosted training or per-move inference services.

The stop rule is about diminishing product returns, not a mathematical ceiling. Future research can reopen with a specific hypothesis and an explicit allowance. For the current release, more unconfirmed model variants are not the main missing feature; a usable and truthful learning interface is.
