One network. Several very different players.

Our Connect Four model ranks moves and estimates positions. Search turns those predictions into a stronger player—but its contribution must be measured separately.

Selected model, local research evaluations · score = (wins + ½ draws) / games
Playing modeCustom ExpertFairy HardGames per opponent
Raw neural policy29.2%19.2%60
Network + 256-simulation MCTS76.5%36%100
Network + native proof/search controller87.5%70%100

Different evaluation banks and sample sizes. These are not paired estimates of the gain from adding each search component, and they are not browser scores.

How it learns

The spatial network has 69,102 learned parameters, seven column logits and a value estimate. It sees each player’s pieces plus engineered immediate-win and threat features. Its training combines self-play positions, certified tactics and policy/value targets from stronger search teachers.

The selected continuation used 4,096 fresh games and repaired historical replay. Unknown outcomes did not receive invented value labels. Immediate wins and unique safe blocks took priority in the teaching targets.

How good is it?

The raw network passed 800 of 800 fresh immediate-win and safe-block tests. That is good tactical evidence, but its raw strategic game scores remain weak. The full local controller supplies much of the playing strength: exact proofs handled about 93.1% of moves in one audit.

Reducing the exact-search cap from five million to one million nodes preserved the declared acceptance criteria and cut exact-search work by 63.9%. The Hard strength gain was statistically uncertain. We selected the cheaper controller without claiming a proven Hard improvement.

Why training stopped

The last policy-teaching candidate scored 77% against Hard, compared with 75% for the accepted model on that round’s fresh bank. Its two-point gain had a paired 95% interval of −4.5 to +8.5 points. It failed the promotion rule, so the accepted checkpoint stayed unchanged.

Three consecutive hypotheses produced no accepted gain. We stopped automatic v1 training and built this learning interface. Further research needs a specific hypothesis and a bounded allowance.

Read our stopping rule

Play the game

The existing Connect Four table lets you practice against the site’s computer opponents or play with another person. The research neural model and its native proof controller are not live website opponents in this release.

Model identity and evidence

Selected spatial model · continuation training seed 193 · one-million-node exact cap · 256 PUCT simulations. Checkpoint SHA-256:

17d089780bc295fcc91b70ee50efa56659e8495c1004dc78d6692b02a310634f

Only the network is exported to ONNX. Browser workers, controller parity and mobile response times remain to be qualified.

Read the complete Connect Four model card