One network. Several very different players.
Our Connect Four model ranks moves and estimates positions. Search turns those predictions into a stronger player—but its contribution must be measured separately.
| Playing mode | Custom Expert | Fairy Hard | Games per opponent |
|---|---|---|---|
| Raw neural policy | 29.2% | 19.2% | 60 |
| Network + 256-simulation MCTS | 76.5% | 36% | 100 |
| Network + native proof/search controller | 87.5% | 70% | 100 |
Different evaluation banks and sample sizes. These are not paired estimates of the gain from adding each search component, and they are not browser scores.
How it learns
The spatial network has 69,102 learned parameters, seven column logits and a value estimate. It sees each player’s pieces plus engineered immediate-win and threat features. Its training combines self-play positions, certified tactics and policy/value targets from stronger search teachers.
The selected continuation used 4,096 fresh games and repaired historical replay. Unknown outcomes did not receive invented value labels. Immediate wins and unique safe blocks took priority in the teaching targets.
How good is it?
The raw network passed 800 of 800 fresh immediate-win and safe-block tests. That is good tactical evidence, but its raw strategic game scores remain weak. The full local controller supplies much of the playing strength: exact proofs handled about 93.1% of moves in one audit.
Reducing the exact-search cap from five million to one million nodes preserved the declared acceptance criteria and cut exact-search work by 63.9%. The Hard strength gain was statistically uncertain. We selected the cheaper controller without claiming a proven Hard improvement.
Why training stopped
The last policy-teaching candidate scored 77% against Hard, compared with 75% for the accepted model on that round’s fresh bank. Its two-point gain had a paired 95% interval of −4.5 to +8.5 points. It failed the promotion rule, so the accepted checkpoint stayed unchanged.
Three consecutive hypotheses produced no accepted gain. We stopped automatic v1 training and built this learning interface. Further research needs a specific hypothesis and a bounded allowance.
Read our stopping rulePlay the game
The existing Connect Four table lets you practice against the site’s computer opponents or play with another person. The research neural model and its native proof controller are not live website opponents in this release.
Model identity and evidence
Selected spatial model · continuation training seed 193 · one-million-node exact cap · 256 PUCT simulations. Checkpoint SHA-256:
17d089780bc295fcc91b70ee50efa56659e8495c1004dc78d6692b02a310634fOnly the network is exported to ONNX. Browser workers, controller parity and mobile response times remain to be qualified.
Read the complete Connect Four model card